Hiring is the highest-stakes repeated decision most organizations make, and it is the one where practice diverges furthest from the evidence. A century of research on what predicts job performance keeps returning the same uncomfortable ranking: the selection methods managers lean on hardest (the free-flowing interview, the résumé read, the practiced gut feel) sit at or near the bottom of the validity table, while the structured, standardized methods they resist sit at the top. Every hire made on a weak predictor is an unpriced bet. The research literature lets you price it.
The evidence base here is unusually deep. In 1998, Frank Schmidt and John Hunter published a synthesis in Psychological Bulletin that compressed 85 years of research findings into a single table: for each common selection method, one number summarizing how well scores at hire predicted job performance later (Schmidt & Hunter, 1998). The table was not a first pass. It extended the ranking of alternative predictors Hunter and Hunter had assembled in 1984, a comparison the field had already spent decades building (Hunter & Hunter, 1984), and it became one of the most influential papers in personnel psychology. Then, in 2022, Paul Sackett and colleagues re-examined the statistical machinery behind it and published revised estimates in the Journal of Applied Psychology: most numbers came down, and the order changed (Sackett, Zhang, Berry, & Lievens, 2022).
A reader could take the 2022 correction as a reason for cynicism: even the flagship numbers moved. The opposite conclusion is the right one. The revision lowered the estimates for nearly every method, including all of the former leaders. That makes the case for structured, multi-method assessment stronger, not weaker, and it leaves the gap between structured measurement and intuition as wide as ever. This article walks through what the evidence says, what a validity coefficient is actually worth in hires, and why organizations keep choosing the methods that predict least.
What predicts job performance, and what merely feels predictive
Start with the unit of measurement. A validity coefficient is a correlation between a predictor (a test score, an interview rating, years of experience) and job performance measured later, usually by supervisor ratings. A coefficient of zero means the predictor tells you nothing: selecting on it is a coin flip. A coefficient of 1.0 would mean perfect foresight, which no method approaches. Selection research expresses every method's value in this common currency, which is what makes a like-for-like ranking possible.
Schmidt and Hunter's 1998 synthesis was a meta-analysis, a study that pools the results of many individual studies to estimate the underlying relationship more precisely than any single study can. Their ranking, reproduced in Figure 1, held the field's consensus for a generation. At the top sat direct, standardized measurement: work sample tests at .54, general mental ability (GMA) tests at .51, structured interviews at .51, job knowledge tests at .48, and integrity tests at .41. In the middle sat unstructured interviews at .38, assessment centers at .37, biographical data at .35, conscientiousness measures at .31, and reference checks at .26.
The bottom of the table is where the discomfort begins, because it is populated by the two staples of every résumé screen. Years of job experience came in at .18. Years of education came in at .10 — barely distinguishable from noise. Graphology, the analysis of handwriting still used in some European markets at the time, scored .02: a coin flip with ceremony. The pattern across all 13 rows is hard to miss. Methods that measure the actual ingredients of performance (can the person do the work, do they know the domain, will they apply themselves) predict well. Methods that infer those ingredients from proxies predict poorly.
Two features of this table deserve a moment before we complicate it. First, the unstructured interview, among the most widely used selection methods, sat at .38, well below the structured version of the same conversation at .51. Same medium, same participants, different discipline, materially different validity. Second, Schmidt and Hunter also reported what happens when valid methods are combined: GMA plus an integrity test reached .65, GMA plus a structured interview .63, and GMA plus a work sample .63. No single method came close to those composites. Hold that thought; it survives everything that follows.
The 2022 correction lowered the numbers and strengthened the argument
Meta-analytic validity estimates are not raw correlations. Studies of selection method validity can only observe people who were actually hired, and the hired are a score-compressed slice of the applicant pool, the phenomenon researchers call range restriction. A correlation computed on that compressed group understates the true relationship in the full pool, so meta-analysts correct the observed figure upward. The size of that correction depends on an assumption: how severely restricted the studied samples were.
This is the assumption Sackett, Zhang, Berry, and Lievens (2022) went after. Re-examining the studies underlying the classic syntheses, they argued that the severe restriction earlier corrections assumed rarely holds in the actual data: many validation samples were not nearly as compressed as the correction formulas presumed. Correcting less aggressively, they produced a revised table: structured interviews at .42, job knowledge tests at .40, empirical biodata at .38, work samples at .33, GMA at .31, integrity tests at .31, assessment centers at .29, situational judgment tests at .26, and unstructured interviews and conscientiousness both at .19.
Figure 2 traces what moved. The steepest falls belonged to the methods most flattered by the old corrections: work samples dropped from .54 to .33 and GMA from .51 to .31. Job knowledge tests held up comparatively well, easing from .48 to .40. Empirical biodata (biographical data scored against statistically validated keys) was the rare climber, from .35 to .38. The structured interview slipped from .51 to .42 but inherited the top of the table. And the unstructured interview was nearly cut in half, from .38 to .19.
It is worth being precise about what did not change, because the reordering can obscure it. Unstructured judgment gained nothing; it was nearly halved. And nothing in the revision rehabilitated the proxy measures at the bottom of the 1998 table; the correction moved the strong methods toward the pack without moving the pack up. The distinctive value of structure survived intact: a structured interview at .42 against an unstructured one at .19 means the discipline of the method, not the conversation itself, carries more than half the predictive weight. With the best single predictor now at .42, no method can carry a hiring decision alone. The 2022 numbers do not soften the case for rigorous assessment; they close the last excuse for relying on any one instrument, however good.
What buys structured interview validity is specifiable, not mystical. In a comprehensive review of interview structure, Campion, Palmer, and Campion (1997) catalogued what the word means in practice: questions derived from an analysis of the job's actual demands, the same questions put to every candidate, answers scored against anchored rating scales, and ratings made independently rather than negotiated in the room. Each element closes a channel through which impression and idiosyncrasy leak into the judgment. An interview is not structured because it feels rigorous; it is structured when those constraints are in place.
Three caveats belong in any honest reading of either table. These are averages across jobs, organizations, and settings; every estimate carries a range around it, and specific roles can sit above or below the mean. The coefficients describe correlations with measured performance, usually supervisor ratings, so the quality of the criterion bounds what any predictor can show. And no coefficient, however high, guarantees an outcome for an individual hire. Validity is about odds, which is exactly why it can be priced.
What a coefficient of .42 buys in actual hires
The most common objection to the evidence on pre-employment assessment validity is that the numbers sound small. A correlation of .42 explains a modest share of variance, and a skeptical manager can round that intuition down to "these tests barely work." The intuition fails because hiring is a repeated binary decision, offer or pass, rather than a variance-explanation exercise. The correct translation of a correlation into that world is an expectancy: given a candidate's assessment score, what are the odds they turn out to be a good performer?
Figure 3 makes the translation with a standard statistical model. Take a candidate who scores in the top fifth of an assessment, and ask the probability that they perform above the median once on the job. With no assessment signal at all (r = 0), that probability is 50% — a coin flip, by definition. At a validity of .10, roughly what years of education delivers, it rises only to 56%. At .30, the neighborhood of the revised estimates for cognitive ability and integrity tests, it reaches 69%, better than two in three. At .50, the neighborhood of the strongest 1998 estimates and of well-built method combinations, it reaches 79%, nearly four in five. A structured-interview-level validity of .42 sits between those last two curves.
The expectancy lens also explains why debates about what predicts job performance so often talk past each other. A researcher reporting a coefficient and a manager asking "will this hire work out?" are answering different questions with the same number. The coefficient promises nothing about one candidate; it sets the odds across many. But organizations do not make one hire; they make a distribution of hires, and the right analogy is underwriting. An insurer cannot say which policyholder will claim; it can still price the book with precision, and it would never accept a process that priced the book at chance. Selection validity prices the hiring book the same way.
Framed this way, the difference between a weak and a strong selection process stops looking academic. Moving from a coin flip to 69% or 79% odds, applied to every offer an organization extends, compounds the way any repeated edge compounds. A firm making hundreds of hires a year on near-zero-validity signals is not merely inefficient; it is systematically stocking its median with people a better process would have screened, and paying for the difference in ramp time, management load, and attrition. Pricing that difference in currency rather than probability is the subject of utility analysis, which we treat in a companion piece on the return on investment of valid selection.
No single method wins; a structured combination does
If the strongest single predictor tops out at .42, the practical question becomes how predictors behave together. The answer is the most durable good news in the literature: selection methods capture partially different slices of performance, so combining them adds signal rather than redundancy. Cognitive measures speak to whether a person can master the work. Integrity and conscientiousness measures speak to whether they will apply themselves to it. Structured interviews and work samples capture applied judgment and interpersonal behavior that neither trait nor ability tests see directly.
Schmidt and Hunter quantified the effect for two-predictor pairs, and Figure 4 shows the pattern. On the 1998 corrections, GMA alone stood at .51; adding an integrity test lifted the composite to .65, and adding a structured interview or a work sample lifted it to .63. Each supplement is individually weaker than the ability test it joins, yet each adds meaningfully, because what it measures overlaps only partially with what the ability test measures. The lesson generalizes: the incremental value of a predictor depends less on its solo ranking than on how little it correlates with what you already measure.
Personality deserves a specific word here, because its broad-trait numbers invite the wrong dismissal. Conscientiousness at .31 in 1998 and .19 in 2022 describes one broad trait predicting one aggregate criterion, overall performance averaged across every kind of job. That is close to the least favorable question you can ask of a personality measure. Which traits matter, and how much, differs by role; and broad traits are themselves bundles of narrower facets that diverge in what they predict. How much resolution a personality instrument should have is a live research question we take up in companion pieces on facet-level personality measurement and on why the Big Five outperforms type-based frameworks. For this article, the composite logic is the point: a valid trait signal that is uncorrelated with ability earns its place in the battery even when its solo coefficient looks modest.
The methods managers trust most are the ones the data trusts least
None of this evidence is new, obscure, or contested at the level that matters for practice. So the persistence of low-validity hiring is itself a finding that needs explaining. Scott Highhouse (2008), writing in Industrial and Organizational Psychology, named the phenomenon directly: a stubborn reliance on intuition and subjectivity in employee selection. Practitioners, his analysis argued, persistently prefer intuitive judgment despite the evidence, and a core barrier is the near-universal belief in one's own skill at reading candidates, a skill the data says almost no one possesses to the degree they believe.
The confidence problem compounds with seniority. The people with final authority over the hiring process are usually those with the longest interviewing histories, and years of practice deepen conviction without necessarily adding accuracy: practice without feedback improves fluency and confidence, not calibration. So the evidence for structure meets its stiffest resistance from precisely the people whose sign-off it needs.
The deeper problem is not the interview itself but the seduction of holistic judgment: the conviction that an experienced human, weighing everything at once, must outperform any formula. This conviction has been tested exhaustively, and it fails. Grove, Zald, Lebow, Snitz, and Nelson (2000) meta-analyzed 136 studies comparing mechanical prediction (combining information by explicit rule) against expert clinical judgment across domains. The formula equaled or outperformed the expert in roughly nine out of ten comparisons; the expert clearly beat the formula in fewer than one in ten. The expert's edge, where it existed at all, was rare enough to be the exception that proves the rule.
Selection has its own version of this result, and it is more pointed. Kuncel, Klieger, Connelly, and Ones (2013) examined what happens when hiring and admissions decisions combine the same information (the identical scores, ratings, and records) either by rule or by holistic judgment. Combining by rule improved predictive validity for job performance by more than half. Read that carefully: the loss is not in what organizations collect, but in the final act of synthesis. A panel that gathers structured, valid evidence and then folds it into a gut-level overall impression is discarding a large fraction of the signal it just paid to acquire.
Why does the belief survive contact with this much disconfirmation? Because hiring gives decision-makers almost no corrective feedback. The rejected candidate's counterfactual career is invisible; the hired candidate's performance arrives slowly, gets attributed to management or luck, and is rarely traced back to the interview that predicted the opposite. An intuition that is never scored cannot lose. That is what makes the empirical literature on what predicts job performance so valuable: it supplies the scoreboard organizations cannot generate for themselves — and the scoreboard shows the formula ahead.
How to run selection as if the evidence mattered
The evidence supports a small number of blunt operating rules. None requires exotic tooling; all require giving up some cherished discretion.
- Structure everything that can be structured. The gap between .42 and .19 measures what discipline adds to the same conversation, not what testing adds over interviewing. Fixed questions tied to the role's demands, anchored rating scales, and ratings recorded independently before any group discussion convert the most trusted ritual in hiring into its strongest single predictor.
- Combine valid, weakly correlated predictors. Pair a can-do measure (ability, job knowledge, work sample) with a will-do measure (integrity, role-relevant traits) and a structured behavioral assessment. The composite logic of Figure 4 (supplements earn their place by measuring what the anchor does not) is the design principle for any assessment battery.
- Measure directly; stop inferring from proxies. Years of experience (.18) and education (.10) are attempts to infer knowledge and ability from biography. Tests of knowledge and ability measure them directly, at higher validity, in less time than a résumé debate consumes. Screen with measurement, and let the résumé provide context rather than verdicts.
- Let the rule make the first cut. Aggregate scores mechanically, weighted and combined the same way for every candidate, and confine human judgment to the stages where it adds information rather than noise: defining the role's demands, probing evidence in structured interviews, and closing candidates. Kuncel's result implies that this one procedural change can recover more validity than upgrading any single instrument.
Two disciplines protect these rules in practice. The first is an evidence trail: record scores, keep the scoring rubric, and periodically check predictor scores against later performance, because published validities are averages and your roles are specific. The second is skepticism at the point of purchase. When evaluating assessment vendors, ask for the validity evidence behind each instrument and how it was corrected: the difference between a 1998-style and a 2022-style estimate is the difference between a marketing number and a planning number. A structured way to run that comparison is laid out in our guide to comparing assessment platforms, and the operational case for administering multiple instruments in a single sitting, which makes multi-predictor measurement practical at hiring volume, is covered in the platform overview.
What the evidence will not do is guarantee any individual outcome. A candidate in the 79% band can fail; one from the 50% band can excel. The claim of valid selection is narrower and more useful: across the repeated decision, the organization that measures directly, combines structurally, and aggregates mechanically will end up with a measurably stronger workforce than the one that trusts its gut — and the size of that edge is no longer a matter of opinion.
Where 5Profiler stands
The evidence in this article reduces to two instructions: measure the ingredients of performance directly, and combine several valid predictors by rule rather than by discussion. That is the design brief 5Profiler was built to. Candidates complete multiple structured instruments covering ability and role-relevant traits, and every candidate's results are scored and weighted identically, so the shortlist comes out of the composite rather than out of whoever spoke last in the debrief. Adaptive testing and role-referenced scoring sit underneath that composite. The structure the literature rewards runs end to end, from item to report, and no step depends on an interviewer's confidence in their own judgment.
References
- Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3), 655–702.
- Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., & Nelson, C. (2000). Clinical versus mechanical prediction: A meta-analysis. Psychological Assessment, 12(1), 19–30.
- Highhouse, S. (2008). Stubborn reliance on intuition and subjectivity in employee selection. Industrial and Organizational Psychology, 1(3), 333–342.
- Hunter, J. E., & Hunter, R. F. (1984). Validity and utility of alternative predictors of job performance. Psychological Bulletin, 96(1), 72–98.
- Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98(6), 1060–1072.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.