A number that organized selection research for a generation moved by .20, and the data beneath it never moved at all. From 1998 onward, the operational validity of general mental ability tests, their estimated correlation with later job performance, stood at .51; in 2022 it was re-estimated at .31. The movement came entirely from the statistical corrections applied to results that were already decades old.
An event like that invites two opposite misreadings. The first treats the revision as a demotion: cognitive testing was oversold, and hiring should look elsewhere. The second treats it as an accounting quarrel among psychometricians, too technical to change any decision. Both readings fail. No party to the dispute argues that ability fails to predict performance; the argument is about how much repair the raw correlations should receive before publication. And because the same repair question touches nearly every selection method at once, "look elsewhere" mostly means looking at methods whose numbers fell too.
For anyone using cognitive ability tests in hiring, the stakes are practical. A vendor quoting a validity of .51 and a rival quoting .31 can both point at peer-reviewed literature, and both will be pointing accurately. Knowing what each number assumes is the difference between trusting a coefficient and trusting measurement.
The .51 was never an observation
Start with what the famous number claimed. A validity coefficient, the correlation between a predictor measured at hire and job performance measured later, is the common currency of selection research; the flagship article in this series ranks every major method in it. The criterion in most studies is a supervisor's rating of overall job performance. And the coefficient printed in meta-analytic tables is something no study directly observed: an estimate of what the correlation would be once known distortions in the underlying studies are repaired.
The first distortion is criterion unreliability: supervisor ratings measure true performance with error, and error in the criterion drags every correlation down, so analysts adjust upward to estimate the relationship with performance itself, not with a noisy rating of it. The second is range restriction, the fact that validation studies can observe only people who were actually hired. Hiring compresses the range of scores, and a correlation computed inside a compressed group understates the correlation that holds in the full applicant pool. Repairing that requires an assumption about how compressed the studied samples were, and the assumption cannot be read off the correlations themselves.
Schmidt and Hunter's 1998 synthesis in Psychological Bulletin, the summary of 85 years of validation research that anchored the field's rankings, applied both corrections and put general mental ability at .51 against overall job performance (Schmidt & Hunter, 1998). Around that anchor the standard advice arranged itself: ability as the backbone predictor, other methods valued partly for what they added over it. How the full table looked, and how it re-sorted in 2022, is the flagship article's territory; this briefing stays on the GMA row.
The point that matters for everything downstream is that .51 was an estimate conditional on a model of typical range restriction, not a fact recorded in any dataset. The corrections were built in from the start, published in the open, and stable for so long that the field stopped seeing them. A generation of textbooks, vendor decks, and utility models carried the number forward while the assumption underneath it rode along unexamined.
The machinery itself was orthodox. Correcting correlations for study artifacts is what meta-analysis exists to do, and rival syntheses of the era all made corrections of their own. What arrived in 2022 was a challenge to the default settings, not to the legitimacy of repair: a claim that the assumed severity of range restriction did not match the samples the field had actually studied.
The .31 rests on a refusal to guess
In the Journal of Applied Psychology, Sackett, Zhang, Berry, and Lievens (2022) went back to the studies behind the classic syntheses and asked whether the assumed compression was actually present. Their conclusion: the severe range restriction that earlier corrections presumed rarely holds in real validation samples. Much of the evidence base comes from concurrent studies, validation run on current employees rather than on applicant pools, and the information needed to justify a large correction was usually missing. Where earlier syntheses corrected aggressively by default, Sackett and colleagues applied less aggressive range-restriction corrections, and general mental ability landed at .31.
The distinction between study designs carries most of the weight in that judgment. A predictive study tests applicants and waits to see how hires perform; a concurrent study tests people already doing the job and correlates scores with current performance. When an estimate comes from a concurrent sample, correction formulas built for applicant pools need a piece of information the study rarely recorded: how the incumbents' score spread compares with the spread of the pool an employer would actually face. Earlier syntheses filled that gap with an assumption of severe restriction. The re-analysis, where the evidence to justify a correction was missing, declined to fill it at all.
The character of that estimate matters as much as its size. The .31 is corrected for criterion unreliability but generally not for range restriction: where solid applicant-pool information was absent, the authors left the correlation unrepaired, a stance they call the principle of conservative estimation. By their own description, that makes .31 a likely slight underestimate. The estimate also comes with spread. The standard deviation of the operational validity distribution is .14, which describes a wide band of true values across settings, not a point. Reading the paper as ".31, give or take a real distribution" is more faithful than quoting any single figure from it.
Figure 1 puts the two estimates side by side, each annotated with what it assumes about restriction. The discomfort the pairing produces is useful. Validity estimates have always traveled through assumptions on their way to print; the revision surfaced uncertainty that was present all along instead of manufacturing it (Sackett et al., 2022). A reader who learned the .51 as a fixed property of ability tests learned something the 1998 authors never quite claimed, and the correction of that habit of reading may outlast the dispute over the coefficient itself.
The rebuttal, the reply, and the interval where GMA validity now lives
The 2022 paper did not settle the question; it opened a live exchange. Oh, Le, and Roth (2023) published a rebuttal in the same journal arguing that the recommendation against correcting in concurrent studies rests on three conditions, that all three fail, and that the revised figures therefore understate operational validity. On their account, the corrections earlier syntheses applied remain appropriate, and the older, higher estimates are closer to the truth (Oh, Le, & Roth, 2023).
Sackett, Berry, Lievens, and Zhang (2023) answered directly. They endorse correcting for range restriction whenever credible information about the applicant pool exists; their position was never that restriction is imaginary. What they maintain is that substantial restriction in concurrent studies is far less common than the rebuttal assumes, and that applying uniform corrections derived from predictive designs systematically overcorrects (Sackett, Berry, Lievens, & Zhang, 2023).
Notice the shape of the disagreement, because it is narrower than the headlines suggested. Both camps agree that ability predicts job performance. Both agree that observed correlations understate true relationships and deserve repair in principle. They differ on an empirical question about archived studies: how often the conditions justifying a large correction were actually met. That bounds the answer. GMA validity for overall job performance sits somewhere between .31 and .51, the revision's own authors call the low end a likely slight underestimate, and the rebuttal's authors argue the shortfall is larger still. What no published party defends is zero, or anything near it.
For practitioners, the exchange rewards a particular kind of reading. Rebuttals of this sort are adversarial about inference and cooperative about facts: the two camps work from the same archived studies, apply the same formulas, and agree on what evidence would settle the matter, namely credible applicant-pool information for the samples in dispute. Where that information exists, they already agree correction is appropriate; the disagreement concerns the default when documentation is missing. The residue for practice is an interval, not a verdict, and intervals are workable; engineering runs on tolerances. What fails is pretending the interval is a point, in either direction. A hiring team can act sensibly at .31, at .51, and at every value between, because the design responses the evidence supports turn out to be the same across the whole interval.
A bounded dispute is an uncomfortable foundation for hiring policy, but it is a normal thing in a working science. The productive response is not to pick a side by temperament but to notice which findings do not depend on the contested assumption at all.
What did not move: general mental ability still scales with job complexity
The complexity gradient heads that list: ability tests predict best where jobs demand the most information processing. The underlying data come from Hunter and Hunter's 1984 analysis of the General Aptitude Test Battery database, a study of aptitude test validity spanning 515 civilian jobs and more than 32,000 employees (Hunter & Hunter, 1984). Schmidt and Hunter's 1998 synthesis relabelled the complexity tiers and reported the corrected validities that Figure 2 plots: .58 for professional-managerial work at the top of the gradient, down to .23 for completely unskilled work at the bottom (Schmidt & Hunter, 1998).
The gradient carries a message the .51-versus-.31 fight cannot touch. The levels in Figure 2 are 1998-era corrected values, so the correction dispute applies to each of them just as it applies to the headline number. What the dispute has no mechanism to undo is the slope: whichever repair policy you prefer, ability predicts most where roles involve novel problems, ambiguous information, and consequential judgment, and least where work is fully proceduralized. The gradient is more useful than any single coefficient, because it says where ability measurement belongs in a hiring design, not merely how much to admire it.
In practice the gradient reads as a weighting rule. For roles built on novel judgment (analysts, engineers, managers, anyone whose day is a sequence of unfamiliar problems), ability deserves a large share of the composite. For heavily proceduralized roles, the same test is a weaker instrument, and other predictors should carry the weight. The gradient also explains why blanket verdicts on testing mislead in both directions: a policy tuned for one tier of complexity exports badly to another, and an organization hiring across tiers should expect its ability weighting to vary by role, not by ideology.
The gradient sits inside a broader regularity. In the 1984 database, ability tests showed validity across all jobs studied: not uniformly high validity, but signal everywhere (Hunter & Hunter, 1984). Two decades on, Schmidt and Hunter reviewed the accumulated evidence on general mental ability in working life and concluded that it predicts both the occupational level a person attains and performance within the chosen occupation, better than any other single trait measured (Schmidt & Hunter, 2004). Cross-occupation generality of that kind is rare in selection research, and nothing in the 2022 revision disturbs it, because generality is a claim about where the signal appears, not about its corrected magnitude.
Training outcomes are a third steady point, with a caveat that belongs in print. The 1998 table put GMA's validity for training performance at .56, above its own figure for job performance (Schmidt & Hunter, 1998). That number is a pre-2022-era corrected estimate: the re-analysis has not yet revisited the ability–training relationship for the same correction issues (Sackett, Zhang, Berry, & Lievens, 2023). Treat it as a strong signal reported under the older, more generous repair policy, not as a value that has already survived the audit. For campus and early-career pipelines, where the job is mostly learning, it remains the first row to consult; the design consequences for graduate hiring are set out on the campus assessment page.
Ability keeps its place in the battery and loses the throne
The sharpest practical consequence of the revision is not the GMA row itself but what happens when ability joins other predictors. Under the 1998 numbers, GMA looked like the natural foundation of any selection composite: pairing it with an integrity test yielded a combined validity of .65, and pairing it with a structured interview or a work sample yielded .63 in either case (Schmidt & Hunter, 1998). On those figures, every serious battery started with an ability test, and the design question was only what to add.
Sackett, Zhang, Berry, and Lievens returned to that logic in 2023 with the revised matrix in hand, in a paper on selection system design in Industrial and Organizational Psychology. Under the prior validity matrix, removing GMA from a six-predictor composite subtracted .20 of validity, a hole no defensible design would accept. Under the revised estimates, the same removal subtracts .05 (Sackett, Zhang, Berry, & Lievens, 2023). Figure 3 shows both subtractions on one scale. Composites with and without ability are now nearly equivalent, and the claim that GMA is irreplaceable no longer has numbers behind it.
The same paper ranks how far each method fell. Against the 1998 table, the largest downward revisions were work samples at −.21, cognitive ability at −.20, and unstructured interviews at −.19 (Sackett, Zhang, Berry, & Lievens, 2023). Ability took one of the deepest cuts, and its relative standing changed accordingly: still a strong predictor, no longer the presumptive core.
Removing a strong predictor takes surprisingly little from a composite, for a structural reason: predictors in a composite share credit. When several instruments capture overlapping signal, the marginal contribution of any one of them is smaller than its standalone validity suggests, and the flatter the revised estimates became, the more the overlap dominates. That mechanism, a property of how composites combine information and not any weakness in the tests, is what moved the design conversation from "ability plus supplements" to "a balanced battery in which ability is one instrument."
Resist the tempting overcorrection here. A gap of .05 is not zero, ability is cheap to measure well, and in complex roles the gradient in Figure 2 argues for weighting it up rather than out. The real lesson of Figure 3 is architectural. A composite that spreads weight across several moderate predictors barely notices when one input is re-estimated; a process built on a single coefficient inherits that coefficient's error bars as its own. There is also a consideration that has nothing to do with validity: cognitive ability carries the largest average subgroup differences of any widely used predictor (Sackett, Zhang, Berry, & Lievens, 2023), and the design implications of that fact get a full treatment in the adverse impact briefing. On both grounds, the multi-method battery is the defensible design, and after the revision it forfeits almost nothing.
Choosing ability measurement while the argument runs
Treat every quoted validity coefficient as an estimate with error bars. The revised GMA figure carries a standard deviation of .14 around it; the older figure carries the contested correction inside it. Across settings, the same instrument will legitimately show different validities, and that spread is information about jobs, not a defect in the test. A coefficient quoted without its correction story is an incomplete sentence, whatever its size.
Weigh measurement quality above any headline number. The correction dispute is about population-level statistics; it says nothing about whether a particular instrument measures ability well, whether its items behave, whether scores stay reliable near the decision point, or whether the item pool resists exposure. Those are properties a vendor either publishes or does not (the measurement evidence behind 5Profiler's instruments is published for exactly this reason). Adaptive administration, which selects each next item at the candidate's estimated level, is one route to that quality; the adaptive testing briefing covers the mechanics.
Run ability inside a battery, not as a sole gate. The revised matrix lowers the stakes of that choice: shifting weight away from ability moves a composite by .05, not .20, so pairing an ability measure with structured interviews and work-relevant assessment sacrifices little and de-risks much (see how the full assessment platform assembles those pieces). Let the complexity gradient set the weight within that battery, heavier where roles run on judgment and lighter where they run on procedure. The battery also answers the subgroup consideration above in the only way that survives scrutiny: by giving strong candidates several routes to demonstrate strength.
Revisit weights and cutoffs when the evidence base moves, because it does. An organization that wrote .51 into its planning models years ago is now running on a disputed input; what a validity point is worth in money is the utility analysis briefing's subject, and the answer scales with the coefficient you believe. Ability remains one of the strongest signals available for the little candidate time it takes to collect, which is exactly what makes recalibration worth the meeting.
The revision strengthens, rather than undercuts, the wider move from pedigree to measurement. The case for replacing degree and school filters with verified assessment, examined in the skills-based hiring briefing, never rested on ability carrying a .5 coefficient; it rests on proxies predicting worse than direct measurement under every correction regime on the table. The interval between .31 and .51 will narrow eventually. A well-designed hiring process does not need to wait for it.
Where 5Profiler stands
5Profiler measures ability with adaptive cognitive assessment and reads the result alongside its other instruments, so a candidate's ability score arrives as one strong signal among several. Each instrument reports separately, which lets an organization re-weight ability as the estimates move, or as its own criterion data accumulates, without rebuilding its process. The same profile carries 30-facet personality resolution next to the ability scores, and weighting can vary by role the way the complexity evidence says it should. That flexibility, held inside a single assessment flow, is the practical response to a coefficient still in dispute.
References
- Hunter, J. E., & Hunter, R. F. (1984). Validity and utility of alternative predictors of job performance. Psychological Bulletin, 96(1), 72–98.
- Oh, I.-S., Le, H., & Roth, P. L. (2023). Revisiting Sackett et al.'s (2022) rationale behind their recommendation against correcting for range restriction in concurrent validation studies. Journal of Applied Psychology, 108(8), 1300–1310.
- Sackett, P. R., Berry, C. M., Lievens, F., & Zhang, C. (2023). Correcting for range restriction in meta-analysis: A reply to Oh et al. (2023). Journal of Applied Psychology, 108(8), 1311–1315.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology: Perspectives on Science and Practice, 16(3), 283–300.
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.
- Schmidt, F. L., & Hunter, J. (2004). General mental ability in the world of work: Occupational attainment and job performance. Journal of Personality and Social Psychology, 86(1), 162–173.