Research Personality science

Same score, different people.

Clustering 1.5 million profiles finds four replicable personality types — as density peaks, not boxes. And scoring the same items at finer grain roughly doubled the variance explained. The case for reading configurations, honestly.

Two assessment reports arrive for one short list printing an identical conscientiousness sten (a band on the standard ten-point scale), an identical percentile, and one shared line of interpretation. Meet the people behind them and the resemblance ends: one is order and caution from end to end, all meticulous plans and careful follow-through with little appetite for stretch goals, while the other is ambition and drive, relentless push and deadlines hit at a sprint with process treated as a suggestion. The instrument has not malfunctioned; both genuinely average out to that one number. But a personality profile is a shape, and the score kept only its height.

This is the difference between a score and a configuration. A score answers "how much"; a configuration answers "arranged how." Most hiring workflows are built entirely on the first question, ranking candidates on one trait at a time as if each trait were a dial that could be read in isolation. But the criterion side of hiring is configural too. Roles do not reward traits one at a time; they reward combinations, and they punish combinations, and two people who match on every marginal score can still differ in the pattern that decides how they will actually work.

An average has many anatomies

The place to start is why the two candidates above can share a score at all. A domain score is an average over narrower facets, and an average is a summary statistic: it preserves the center of a set of numbers and discards their arrangement. High-plus-low equals medium-plus-medium. A sten of 7 can be assembled from order without drive, drive without order, or a moderate helping of both. Nothing in the reported number distinguishes these anatomies, because the averaging happened before the report was printed.

The mapping only runs one way. From a facet profile you can always compute the domain score; from the domain score you cannot recover the profile, because many distinct profiles compress to one value. At the scale of a hiring pipeline that many-to-one compression has a concrete meaning: every point on the score scale is a bucket holding genuinely different people, and one of those buckets holds both candidates this article opened with. The compression is not visible from inside the report, either, which is what makes it dangerous. A domain score carries no marker distinguishing the profile that was uniform from the profile that was jagged, so a reader has no way to know which kind of number is in front of them. Downstream decisions inherit the ambiguity. An interviewer briefed on "high conscientiousness" will probe for reliability and find it in both candidates, because both can truthfully perform the summary. The questions that would separate them live at the level the report never showed.

The measurement case for descending below the domain average, the bandwidth-fidelity question of matching a measure's grain to a decision's grain, is made in full in our companion briefing on facet-level personality measurement; this article takes the next step. Suppose you have the facets. The question becomes what to do with a whole pattern of them, and here the industry splits into two bad habits. The first collapses the pattern back into summaries and ranks candidates domain by domain, which recreates the original problem one level up. The second goes the other way and sorts people into named types, trading the pattern for a box. The research record has something to say about both.

The modern case for personality types comes with a warning label

Typing people is the older instinct. Long before trait scores existed, observers grouped people into kinds, and the commercial assessment market still runs heavily on that instinct: a handful of letters, a grid of boxes, a color, a label. The scientific critique of discrete type tests, including their reliability at the cut points where one type flips into another, is the territory of our review of the Big Five versus personality type indicators, and it is not repeated here. The more interesting question is what happens when a serious research team goes looking for types in trait data with modern methods and modern sample sizes.

That test arrived when Gerlach, Farb, Revelle, and Amaral (2018) published a cluster analysis of four datasets totaling more than 1.5 million respondents in Nature Human Behaviour. Earlier attempts at empirically derived types had typically used samples under about 1,000 people, and their solutions failed to reproduce from study to study. Gerlach and colleagues had the scale to do it properly, and they found something real: at least four replicable personality types, which they named average, reserved, role model, and self-centered. The role model type combines low neuroticism with high scores on the other four domains; the average type sits within roughly one standard deviation of the origin on every dimension.

Then come the qualifications, and they are not fine print; they are the finding. First, the types are density peaks, not boxes. Figure 1 sketches the geometry: people fill the entire trait space, and a type is a region where more of them gather than chance would predict. There are no boundaries to fall on either side of, no membership to test positive for. Calling someone a "role model type" says only which hill of the probability terrain they stand nearest.

Second, the machinery that finds clusters will find them in almost anything. The solution that fit the data best by the standard statistical criterion contained 13 clusters. When the team applied a stricter test, keeping only clusters where real respondents concentrated more densely than in randomized versions of that data, nine of the 13 evaporated. Most of what a standard clustering method confidently reports, in other words, is structure the method invented. Any vendor who has run k-means over an applicant database and announced the discovery of a tidy set of "candidate personas" has performed the first step of this analysis and skipped the step that mattered.

Third, the types are only as recoverable as the measurement is fine. All four types were recovered in the dataset built on a 120-item instrument; the datasets built on 100-item and 44-item measures recovered fewer. The person-space a method can resolve depends on the grain of the items feeding it, a point this article returns to, because it connects the type literature to the resolution literature. Even the study's own replication was partial: an identical analysis, pointed at coarser measurements, could no longer see all of the structure. Anyone deploying type language on the back of a brief screener is claiming resolution the underlying measurement cannot deliver.

The Gerlach study also echoes an older result. A tradition running through developmental psychology repeatedly recovered three prototypes, labeled resilient, overcontrolled, and undercontrolled, in both children and adults, with undercontrollers showing externalizing correlates and overcontrollers internalizing ones (Asendorpf, Borkenau, Ostendorf, & van Aken, 2001). That tradition, too, spent years in replication controversy: the prototypes emerged in some samples and methods and not others, and the field argued at length about whether they were robust (Donnellan & Robins, 2010). The pattern across both generations of research is consistent. Some configural structure in personality is real. It is far less crisp, far less categorical, and far easier to fake with bad methods than the type-test market implies.

The same items, scored at finer grain, roughly doubled prediction

If types over-compress, the natural question is whether de-compressing pays. Mershon and Gorsuch (1988) answered it with a study designed to isolate resolution itself. Across 16 datasets, they scored the same personality item pool two ways, once as six broad factors and once as 16 narrower ones, and asked which scoring predicted real-life criteria better. Nothing else varied: same respondents, same items, same outcomes. The only difference was how much configuration the scoring preserved.

The finer scoring won, and not narrowly. Gains in multiple R (the correlation between a set of predictors taken together and an outcome) were statistically significant in 88% of the comparisons, and the variance accounted for in the real-life criteria roughly doubled even after shrinkage, the adjustment that removes the advantage extra predictors gain by fitting noise. Figure 2 indexes the result: whatever the six broad scores explained, the 16 finer scores explained about twice as much. The information was in the items all along. The coarse scoring was throwing it away, exactly as an average must.

The counterargument comes from serious people, and it is partly right. Ones and Viswesvaran (1996) argued that for predicting broad criteria, above all overall job performance, broad measures are the appropriate bandwidth: a wide criterion is best matched by a wide predictor, and some of the apparent advantage of narrow traits reflects statistical artifacts of how the comparisons are run. As a caution against maximal splitting, the point stands. Nobody should score hundreds of micro-scales and stepwise-regress their way to a spuriously impressive fit.

But read carefully, the two positions converge on one principle: match the grain of the measurement to the grain of the question. When the question is broad, a broad score serves. When the question is a specific role with specific failure modes, the question itself is configural, and a summary is the wrong grain. Most hiring questions are of the second kind, which is what makes the convergence practically consequential. An organization is rarely trying to forecast overall job performance in the abstract; it is trying to decide whether this candidate will hold up in this job, where the failure modes are known and specific. It is worth remembering how modest the broad scores are even on their home turf: under the revised meta-analytic estimates, conscientiousness predicts overall job performance at .19 (Sackett, Zhang, Berry, & Lievens, 2022). Domain scores are not a fortress that configuration-reading proposes to abandon. They are a floor that finer reading has room to build on.

The wrong way to compare a personality profile to a role

Suppose an organization accepts all of this and starts working with whole configurations. There is a disciplined way to do that and a treacherous one, and the treacherous one is more common because it produces a tidier number. The disciplined tradition is person-centered assessment in the methodological sense: latent profile analysis and related mixture methods, which identify subgroups of people sharing a configuration instead of ranking everyone on one dimension at a time (Meyer & Morin, 2016). Meyer and Morin built their methodological program for organizational research, and its logic transfers directly to assessment: the analysis asks which configurations of standing actually occur in a population, and who exhibits them, before any ranking begins. The safeguards are the point. Rival solutions are compared formally, and a profile is retained only when the evidence supports it. How appealing a cluster's story sounds carries no weight in that decision. Everything Gerlach and colleagues did is this tradition operating at scale, null models and all. Practiced that way, person-centered methods are how configurations enter research legitimately.

The treacherous path is the profile similarity index: compute the gap between a candidate's personality profile and a role's target profile on each dimension, collapse the gaps into one number, and rank candidates by it. Edwards (1993) took this practice apart in Personnel Psychology, and the critique has never been answered. A single similarity index is conceptually ambiguous, because very different patterns of disagreement produce indistinguishable values. It conceals which side of the comparison drives the difference and in which direction. It conflates the effects of its constituent dimensions, so nobody can tell which facet's mismatch actually relates to the outcome. And it imposes restrictive constraints that the data are never asked whether they satisfy, such as treating overshooting a target and undershooting it as equivalent errors. His prescription was to keep both profiles intact and model them jointly, with polynomial regression rather than a collapsed score, so the data can reveal which components matter and how.

Figure 3 stages the problem at hiring scale. Candidate 1 matches the role's target on nearly everything and collapses on a single facet, the profile of a specific, nameable risk. Candidate 2 drifts a little on several facets, high on some and low on others, the ordinary texture of a plausible hire. Any similarity index built by summing gaps declares them interchangeable. A recruiter who sees only the fit score has been told two different stories in a single word, and the moment a threshold is applied to that score, the difference between the stories becomes a hiring decision.

Labels can caption a personality profile; they cannot replace it

What should change in reports, then, is less about adding information and more about refusing to destroy it. A report built for personality configurations shows the facet profile itself, read against the demands of the role, with each gap visible in size and direction. Summaries and labels can sit on top of that display as entry points; the failure mode is when they become the display. You can see the difference in practice in a sample assessment report: the domain-level story appears, but every domain opens into the facets beneath it, and the role comparison never collapses into a single number.

Archetype labels deserve what the evidence gives types: real utility, strict supervision. A label is a communication device, a way for a panel to hold a configuration in mind, and labels anchored to measured profiles genuinely help non-specialist readers navigate. The Gerlach results even license a modest scientific status for a few of them, as names for density peaks. What no result licenses is the label as measurement, the report that says "Analyst type" where the profile should be. That substitution is how popular constructs run ahead of their evidence: an evocative name, a tidy box, and the measured configuration quietly retired from the page.

Anchoring is an operational property with testable conditions. A label is anchored when it is computed from the measured profile it names, when a reader can move from the label to the facets in one step, and when two people who share the label demonstrably share the configuration. It is unanchored when the label arrives first and the measurement is fitted to it afterward, the ordering that has produced typologies with better marketing than psychometrics.

The instrument-length finding closes the loop on measurement itself. Gerlach and colleagues recovered all four types only in their 120-item dataset; the shorter instruments recovered fewer. Resolution in the person-space is bought with resolution in the item pool, and no analytic cleverness downstream can recover structure the instrument never captured. An assessment that administers a handful of items per domain cannot support configural reading, whatever its reporting layer claims; the measurement foundations that make facet-level resolution defensible are described on our science page. This is Figure 2's lesson arriving from the opposite direction: grain is where the information lives.

What this means for how organizations use personality assessment

The practical program follows from the three results, and it requires no research department. First, stop gating on domain composites when the facets beneath them disagree. A composite is trustworthy as a decision variable only when it summarizes a reasonably uniform profile; when facets diverge, the honest move is to read them, not to average the disagreement away. The two candidates from the opening paragraph should generate different conversations, and any workflow in which they cannot is discarding measured information. In practice this is a report-design requirement as much as a policy: a composite whose facets disagree should say so on its face, and the disagreement should become interview questions about the specific facets that diverge, which is cheaper than discovering the divergence during the first quarter on the job.

Second, a type or persona label is a compression, and the question is what it compresses. The underlying continuous scores exist; a credible provider can always show them, and the Gerlach findings supply the questions that separate discovered structure from manufactured structure. Was the cluster solution tested against a randomized null? Does it replicate across samples? Is membership presented as a probability or a verdict? Personality types in hiring are not automatically pseudoscience, but the burden of proof sits with whoever drew the boundaries.

Third, the candidate-to-role comparison has to keep both profiles intact. That means facet-level gaps reported with their direction, one large mismatch distinguished from many small ones, and no one-number fit score of the kind Figure 3 shows treating two different candidates as twins. Edwards's alternative is a modeling standard for researchers, but its spirit translates directly into reporting: never let a similarity index answer a question that requires seeing the two profiles it consumed.

These practices do not ask organizations to abandon scores. Scores are how measurement travels, and modest validity, plainly stated, is still value, and the revised estimates leave even the best-performing domain a long way short of a decisive signal. The point is narrower and more actionable. Between the score and the decision sits a layer of information, the configuration, that current reporting habits routinely delete. Organizations that keep it get sharper interviews, more specific onboarding, and fewer surprises of the kind the opening pair was built to illustrate. The two candidates were never the same person. The report should have said so.

Where 5Profiler stands

5Profiler treats the configuration as the deliverable. Every domain score in a report opens into its facets, and role-referenced scoring keeps the candidate-to-role comparison at facet grain: each gap is shown with its size and its direction, and no single fit number is asked to stand in for the pattern. Where a report uses an archetype label, the label is anchored to the measured profile it summarizes, and the profile stays on the page. When the facets inside a domain disagree, the report says so at the point the domain score appears and names the facets that diverge, so a panel walks into the interview knowing which specific questions the summary left unanswered.

Read the science behind the platform · See it on your roles

References

  1. Asendorpf, J. B., Borkenau, P., Ostendorf, F., & van Aken, M. A. G. (2001). Carving personality description at its joints: Confirmation of three replicable personality prototypes for both children and adults. European Journal of Personality, 15(3), 169–198.
  2. Donnellan, M. B., & Robins, R. W. (2010). Resilient, overcontrolled, and undercontrolled personality types: Issues and controversies. Social and Personality Psychology Compass, 4(11), 1070–1083.
  3. Edwards, J. R. (1993). Problems with the use of profile similarity indices in the study of congruence in organizational research. Personnel Psychology, 46(3), 641–665.
  4. Gerlach, M., Farb, B., Revelle, W., & Amaral, L. A. N. (2018). A robust data-driven approach identifies four personality types across four large data sets. Nature Human Behaviour, 2(10), 735–742.
  5. Mershon, B., & Gorsuch, R. L. (1988). Number of factors in the personality sphere: Does increase in factors increase predictability of real-life criteria? Journal of Personality and Social Psychology, 55(4), 675–680.
  6. Meyer, J. P., & Morin, A. J. S. (2016). A person-centered approach to commitment research: Theory, research, and methodology. Journal of Organizational Behavior, 37(4), 584–612.
  7. Ones, D. S., & Viswesvaran, C. (1996). Bandwidth-fidelity dilemma in personality measurement for personnel selection. Journal of Organizational Behavior, 17(6), 609–626.
  8. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.