Somewhere in your organization this week, a hiring debrief will cite a four-letter type. The badge sits in email signatures and team wikis, and, less openly, in judgments about who fits a role. Set against that confidence is one of the most consistently replicated findings in the assessment literature: when people retake a type test after a short interval, between 39% and 76% of them come back as a different type, at intervals as short as five weeks (Pittenger, 1993). The Big Five vs MBTI question, framed honestly, is not a contest between rival personality questionnaires. It is a choice between a measurement tradition and a conversation tool, and only one of the two was built to carry the weight of an employment decision.
That verdict needs a careful defense, because type indicators are not fringe products and their users are not fools. The instruments are embedded in leadership programs, onboarding curricula, and team rituals at many sophisticated companies, and the experience of taking one is genuinely engaging. The case against them in selection is therefore worth building slowly: from the design decision at their core, through the retest record, to the published evidence on predicting work outcomes. It ends somewhere more interesting than mockery. Type frameworks fail in hiring for the same reasons they succeed in workshops.
The Big Five, by contrast, is not a product but a research tradition: a five-dimensional description of personality that emerged from decades of statistical work on how trait descriptions covary, that replicates across languages and cultures (McCrae & Costa, 1997), and that anchors the meta-analytic literature connecting personality to job performance, leadership, and consequential life outcomes. The Big Five vs MBTI comparison that follows is asymmetric because the evidence is asymmetric.
Types came from a theory; traits came from the data
The type indicator descends from Carl Jung's early-twentieth-century speculation about psychological opposites (introversion and extraversion, sensing and intuition, thinking and feeling), proposed as a clinical framework long before modern psychometrics existed. Jung was theorizing, not measuring. The questionnaire that later operationalized his typology, developed by Katharine Briggs and Isabel Briggs Myers in the mid-twentieth century, added a fourth pair, judging versus perceiving, and a scoring rule that assigns every respondent to one pole of each. Four binary choices yield 16 types, and the 16 types yield the badge.
Everything that follows psychometrically is downstream of one commitment: that people come in kinds. A type system does not claim you are somewhat more extraverted than average. It claims you are an extravert — a member of a discrete category, different in nature from an introvert, the way a left-handed person differs in kind from a right-handed one. That is a strong empirical claim, and it is checkable, because categories of people should leave categorical marks in the data.
The five-factor model traveled in the opposite direction. Instead of starting from a theory and building an instrument to confirm it, researchers started from the trait vocabulary of natural language and applied factor analysis (a statistical method for finding the smaller set of independent dimensions underlying many correlated measurements) to ask how many dimensions personality description actually needs. The answer, replicated across languages and cultures, was five (McCrae & Costa, 1997). The model earned its standing not by being appealing but by refusing to go away, and the documentation standards that grew up around it, from published reliability to validity evidence and norms, are the backbone of the science behind modern trait assessment.
The distinction matters commercially as well as scientifically. A type indicator is a product: one publisher, one questionnaire, one scoring key. The five-factor model is a literature: many competing instruments measure it, independent research groups test it, and no company controls what it is permitted to say. When a framework's claims are policed by rivals rather than by its owner, the claims that survive mean more.
The strangest fact in this literature is friendly to the type camp. When McCrae and Costa (1989) correlated the indicator's continuous scores with five-factor measures, they found the MBTI scales track four of the five major trait dimensions. The item pool is not astrology; it captures real trait variance. What the type system then does with that variance, cutting each continuous score in half, is where the trouble begins.
The cut falls exactly where the people are
A theory of kinds makes a testable prediction. If the working population divides into introverts and extraverts the way it divides into left- and right-handers, preference scores should be bimodal: two humps with a valley between them, and few people near the boundary. That is not what the data show. McCrae and Costa (1989) found no evidence of bimodality in the indicator's continuous scores. The distributions are unimodal, with respondents piling up near the middle of each scale, which is precisely where the type cut is drawn.
Every psychological score carries measurement error, and psychometricians quantify it as the standard error of measurement: the expected wobble in a person's score from one administration to the next, driven by item sampling, mood, and context rather than any real change. For a respondent near the middle of a preference distribution (which is to say, for most respondents), routine wobble is enough to cross the cut. Nothing about the person changes; the letter does. Figure 1 shows the geometry of the problem.
Dichotomization also discards magnitude. Two colleagues a few points on either side of the boundary, psychometric near-twins, receive opposite letters and opposite narratives, while two people at opposite extremes of the same half share a label and a narrative. The type report is most confident about exactly the people the measurement can least distinguish, and least informative about the differences it could measure well.
None of this depends on bad items or careless administration. It follows mechanically from cutting a continuous, unimodal distribution at its densest point. A type indicator with flawless items would still reclassify a large share of its takers on a second sitting, which is precisely what the retest literature shows.
An instrument that changes its mind within five weeks
The prediction implicit in Figure 1 has been tested directly. Reviewing the retest studies, Pittenger (1993) found that between 39% and 76% of respondents were assigned a different four-letter type when they retook the indicator, in some studies after a gap of only five weeks. Read the range at its most charitable end and the conclusion is still disqualifying: an instrument consulted in decisions about people hands roughly two in five of them a different identity a few weeks later. At the other end, it is three in four.
Any measure moves on retest; the question is how it fails. A continuous score that drifts a few points degrades gracefully — the profile shifts slightly, the interpretation barely. A category flip is discontinuous. A change of one letter replaces the description wholesale, with different advice attached and, in a hiring context, a different verdict. The format converts ordinary measurement noise into wholesale reclassification.
The practical translation is stark. If an offer, a promotion, or a place on a shortlist is influenced by the letters, then the outcome depends in part on which week the candidate happened to sit the assessment. No hiring process would tolerate a reference check that reversed itself on a second phone call; a classification that reverses for two in five takers (and sometimes three in four) deserves the same scrutiny.
Trait instruments built for measurement handle error differently: they publish it. The professional manual for the NEO PI-R, the reference inventory of the five-factor tradition, reports internal consistencies (the degree to which a scale's items agree with one another) between .86 and .92 for its domain scales (Costa & McCrae, 1992). A buyer can read the error before trusting the score, which is what distinguishes an instrument from a product.
The deeper contrast is longitudinal. Roberts and DelVecchio (2000), in a quantitative review of longitudinal studies published in Psychological Bulletin, tracked the rank-order consistency of traits: whether people hold their standing relative to others as time passes. Consistency rises steadily across the lifespan: roughly .31 in childhood, .54 in the college years, .64 around age 30, plateauing near .74 between ages 50 and 70. People change, but they change in an orderly way that preserves comparison; stability is highest across exactly the ages most of the workforce occupies. Figure 2 sets the two records side by side.
The two panels are not measuring the same thing (the left is a categorical reassignment rate, the right a correlation), and that mismatch is itself the finding. A trait framework can report its stability on a continuous scale and improve interpretation accordingly. A type label assigned near the cut cannot be stable, however good the items, because stability of the label would require the score to stop behaving like a measurement.
Big Five vs MBTI on the only question hiring asks
Selection has one criterion that outranks the others: does the score predict how people will actually perform? The field summarizes this as a validity coefficient, the correlation between scores at hire and outcomes later. Any personality test for hiring stands or falls on that number and on the documentation behind it.
The modern record begins with Barrick and Mount (1991), a meta-analysis (a statistical pooling of prior studies) published in Personnel Psychology that has organized the field since. Its signature finding was not a large effect but a consistent one: conscientiousness predicted job performance in every occupational group studied, with a corrected correlation of about .22. No type category has ever produced a comparable published record (National Research Council, 1991).
Three decades later, Sackett, Zhang, Berry, and Lievens (2022) re-estimated the selection literature in the Journal of Applied Psychology, applying less aggressive range-restriction corrections, and placed conscientiousness near .19. The plain headline is that single-trait validities are modest, well below structured interviews at .42, and a vendor claiming otherwise is overselling. But cheap, stackable, low-friction signals are the currency of good selection systems: as our companion review of what actually predicts job performance details, no single predictor is decisive, and the gains come from combining structured methods rather than betting on one oracle.
Personality's reach also extends beyond the single job. Judge, Bono, Ilies, and Gerhardt (2002) meta-analyzed personality and leadership and found the five traits jointly correlate .48 with leadership (a multiple correlation, meaning the five dimensions working together), with extraversion the strongest single dimension at .31. And in a broad review of consequential outcomes, Ozer and Benet-Martínez (2006) documented trait effects on health, relationship quality, and occupational attainment. Whatever the Big Five's limits, it is connected to the world it claims to describe.
Now the empty row of Figure 3. When a National Research Council committee reviewed techniques for enhancing human performance, it concluded there was no adequate research basis for using the MBTI in performance contexts (National Research Council, 1991). Fourteen years later, Pittenger's (2005) survey of the accumulated psychometric evidence reached the same cautionary verdict on workplace decision use. Most telling of all, the instrument's own publisher advises that the indicator is not designed for, and should not be used for, hiring or selection (The Myers-Briggs Company, n.d.). The gap in the chart is not waiting to be filled by research. It marks a use the instrument's makers themselves disclaim.
This is the asymmetry that settles the Big Five vs MBTI question for selection. One tradition publishes effect sizes that can be criticized, re-estimated, and, as Sackett and colleagues showed, revised downward in public. The other offers no adequate estimate to criticize. In measurement, the capacity to be wrong in print is not a weakness; it is the qualification for being used.
The steelman: types are popular for good reasons
It would be convenient to conclude that type frameworks persist through ignorance. The truth is less flattering to the critics: types deliver real value, and what they deliver is exactly what trait reports have historically delivered badly.
First, a shared vocabulary. A team that has been through a type workshop can say "you want the full picture before deciding; I decide fast and revise" without anyone standing accused of a character flaw. Conversations about personality types at work give colleagues a script for differences that would otherwise surface as friction, and the script is learnable in an afternoon.
Second, safety by design. Every type is written as a strengths profile; there is no bad four-letter combination and no low score, because there are no scores at all in the report the taker sees. That is why the format works in a room: nobody leaves labeled deficient, so everybody participates. A categorical identity is also memorable in a way a vector of five numbers is not — people carry their letters for years, which keeps the vocabulary alive long after the workshop ends.
Third, the indicator, used as its publisher intends, is a prompt for self-reflection rather than a verdict on anyone. Even Pittenger's (2005) critique is aimed at decision use, not at conversation. In a development offsite, the cost of misclassification is an hour of mild misdescription, cheerfully corrected over coffee. The failure modes documented above only become consequential when someone's livelihood depends on the output.
Where workshop virtues become selection liabilities
Each virtue inverts under selection pressure. Safety by design means no candidate can score poorly — which means the instrument cannot rank, and ranking is what selection is. An assessment engineered so that every outcome flatters has, by construction, removed the information a hiring decision requires. The workshop feature is the screening defect.
Comparability fails next. Categories collapse magnitude: two candidates sharing four letters may sit at opposite ends of every underlying scale, yet the report tells the panel they are the same kind of person, while a third candidate a hair across one boundary is presented as categorically different. Continuous trait scores preserve exactly the comparative information the letters throw away, and comparison between candidates is the entire task.
Then there is the audit. A selection procedure should survive one question: what is the evidence that this predicts performance, and how stable is the score you acted on? For a documented trait-based assessment, the answer is a technical manual and a meta-analytic literature. For a type indicator, the answer is a retest record that reclassifies a large share of takers within weeks, a national-academy review that found no adequate research basis, and a publisher that disclaims the use in writing. It is hard to construct a weaker position from which to defend a contested hiring or promotion decision, psychometrically or legally.
This is also what makes the popular jibe that the badge is corporate astrology half wrong and half right. It is wrong about the items: unlike star signs, the indicator's continuous scores track four of the five major trait dimensions (McCrae & Costa, 1989), which is to say they measure something real. It is right about the label: the type layered on top adds a category where the distributions show none, and an unstable category at that. The liability is not the questionnaire. It is the letters.
What this means for buying and running assessment
None of this argues for banishing personality from hiring. The evidence argues for more personality measurement, done to a higher standard — the standard the trait tradition already publishes. Four rules follow from the record.
- Buy instruments built and documented for selection. A serious instrument publishes its error and its evidence, the way the NEO tradition's manual reports domain internal consistencies of .86 to .92 (Costa & McCrae, 1992), and a serious vendor hands over the documentation on request. If the technical manual does not exist, or the publisher itself disclaims selection use, the evaluation is already over.
- Insist on continuous scores read against role demands. A trait score means little in the abstract; the same standing on a dimension can be an asset in one role and a liability in another. Demand norm-referenced, role-specific interpretation rather than universal archetypes. Domain scores are still coarse, and the case for going a level deeper is made in our companion piece on facet-level personality measurement.
- Treat personality as one signal in a structured system. At validities near .19 (Sackett et al., 2022) to .22 (Barrick & Mount, 1991) for conscientiousness, against .42 for structured interviews (Sackett et al., 2022), no personality test for hiring should gate a decision alone. Its role is incremental: an inexpensive, scalable signal inside a documented combination of predictors.
- Give types a firewall, not a funeral. If the four-letter workshop is valuable to your teams, keep it on the development side, where its virtues operate and its instability is harmless. Then write the boundary into assessment governance: no categorical instrument in screening, ranking, or promotion, ever.
The simplest test of any report you are handed is whether you could explain the score to the candidate, or to a regulator, and point to the evidence behind each claim; a fully documented trait report shows what that looks like in practice. Framed this way, the choice between the Big Five and personality types stops being a matter of taste. One option is a measurement tradition with published error bars; the other is a conversation format with a published disclaimer.
Where 5Profiler stands
5Profiler was built as a measurement instrument rather than a conversation tool. Candidates are assessed in the Big Five trait tradition on continuous, norm-referenced scales, never assigned type labels, and the measurement model behind every score is documented so that a review panel, or a skeptical regulator, has something concrete to interrogate: how the scales are built, what they measure, and how reliably. Where a type badge collapses candidates into 16 boxes, the platform resolves personality at 30-facet resolution, preserving the comparative detail a selection decision turns on. Nothing in a 5Profiler report asks a hiring team to trust an archetype; the case for each score is laid out to be questioned, and to survive it.
References
- Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26.
- Costa, P. T., & McCrae, R. R. (1992). NEO PI-R professional manual. Odessa, FL: Psychological Assessment Resources.
- Judge, T. A., Bono, J. E., Ilies, R., & Gerhardt, M. W. (2002). Personality and leadership: A qualitative and quantitative review. Journal of Applied Psychology, 87(4), 765–780.
- McCrae, R. R., & Costa, P. T. (1989). Reinterpreting the Myers-Briggs Type Indicator: Implications for the five-factor model of personality. Journal of Personality, 57(1), 17–40.
- McCrae, R. R., & Costa, P. T. (1997). Personality trait structure as a human universal. American Psychologist, 52(5), 509–516.
- National Research Council (1991). Druckman, D., & Bjork, R. A. (Eds.), In the Mind's Eye: Enhancing Human Performance. Washington, DC: National Academy Press.
- Ozer, D. J., & Benet-Martínez, V. (2006). Personality and the prediction of consequential outcomes. Annual Review of Psychology, 57, 401–421.
- Pittenger, D. J. (1993). Measuring the MBTI… and coming up short. Journal of Career Planning and Employment, 54(1), 48–52.
- Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221.
- Roberts, B. W., & DelVecchio, W. F. (2000). The rank-order consistency of personality traits from childhood to old age: A quantitative review of longitudinal studies. Psychological Bulletin, 126(1), 3–25.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- The Myers-Briggs Company (n.d.). Ethical use guidelines for the MBTI instrument.