Research Cost of getting it wrong

Résumés are claims, not evidence.

Years of education predict performance at .10, experience at .18 — and identical résumés draw 50% more callbacks with a different name at the top. The case for screening on measurement instead of paper.

A recruiter working a high-volume requisition gives each document a few seconds. At that rate the eye does not read; it pattern-matches, catching a familiar employer, a degree line, a job title, the shape of the page, the name at the top. Whatever survives that pass becomes the shortlist, and every rigorous stage that follows, the structured interviews and the assessments and the reference checks, operates only on the survivors. Résumé screening is the one stage of hiring that nearly every organization runs and almost none has ever validated. The instrument deciding who reaches the valid instruments is itself graded by no one.

The research on what résumés predict is not ambiguous, and it has been available for decades. The two lines a screener studies hardest, education and experience, correlate with later job performance at .10 and .18, close to noise. The version of biographical information that genuinely predicts performance is precisely the version no human reader can apply by eye. And the signal that travels most reliably off the page is not job-relevant at all: in a field experiment described below, the same résumé drew roughly 50% more interview callbacks when the name at the top sounded white.

This article works through that evidence, from the validity of the résumé's individual lines to the bias the document carries and the automation that industrialized both. The conclusion is architectural rather than technical. The fix is not a better résumé-reading method, because reading is the problem. It is to move measurement to the front of the funnel, screen every applicant with a short, valid assessment, and demote the résumé to what it actually is: a set of claims to verify before an offer.

The two lines everyone reads are close to noise

Ask what a résumé screen is for and the answer is some version of the same belief: education and experience tell you who can do the job. That belief has been tested against data for longer than most companies have existed. In Schmidt and Hunter's 1998 synthesis of the selection literature, the ranking examined at length in our companion article on what predicts job performance, years of education correlated .10 with later job performance and years of job experience correlated .18 (Schmidt & Hunter, 1998). Those validity coefficients, correlations between a predictor at hire and performance measured later, place the résumé's two most scrutinized lines barely above the floor of a scale on which zero means no information at all.

Set that against what the same literature reports for instruments designed to predict. The strongest structured measurement methods sat at .5 and above in the 1998 synthesis. The 2022 re-analysis by Sackett, Zhang, Berry, and Lievens, which applied less aggressive corrections for range restriction and lowered the estimates for most methods, still places the structured interview at .42 and empirically keyed biographical data at .38 (Sackett et al., 2022). Figure 1 draws the comparison directly. The lines organizations treat as the entry gate carry a fraction of the signal of the methods they postpone until the shortlist.

It is worth pausing on why the staple lines are weak, because the weakness is structural, not accidental. Years of experience counts years — not what filled them. A decade of deliberate, well-coached practice and a decade of repeating the same first year both print as the same integer. A degree certifies that a person cleared an admissions bar and persisted through a curriculum, both of which relate to ability and diligence, but it certifies them at a distance of years and through institutions of wildly uneven meaning. Proxies compress; compression destroys signal. Methods that measure ability, knowledge, or role-relevant behavior directly outpredict biography because they measure the thing itself rather than its residue.

A screener might concede the point and still shrug: screening is cheap, and a weak filter is surely better than none. But a screen is not one decision. It is thousands of small verdicts, and a near-noise instrument applied at that volume produces errors in both directions, silently. Strong candidates cut at the top of the funnel are gone before any valid method can see them, and weak candidates advanced by an impressive-looking page consume interview capacity or become hires. What one of those mistakes costs, in replacement, ramp time, and team drag, is the subject of our companion piece on the cost of a bad hire. Résumé screening is where those mistakes are manufactured in bulk.

The strongest signal on the page is the name at the top

If the résumé's content is nearly uninformative, the next question is what a résumé screen responds to instead. Labor economics supplied the answer with a design built to remove the explanations that contaminate ordinary comparisons: the audit study. Researchers construct fictitious, matched applications, vary a single attribute at random, send them to real job openings, and record which ones receive a callback, an invitation to interview. Because the attribute is assigned by lottery, no difference in qualifications, motivation, or presentation can account for a difference in outcomes. Whatever gap appears was produced by employers' response to the manipulated attribute and nothing else.

Marianne Bertrand and Sendhil Mullainathan ran that design at scale. In a field experiment published in the American Economic Review, they answered help-wanted ads in Boston and Chicago with 4,870 fictitious résumés, assigning each a name that reads as white, like Emily or Greg, or as African-American, like Lakisha or Jamal (Bertrand & Mullainathan, 2004). The content under the name did not differ systematically; the name was the experiment. Résumés with white-sounding names received a callback 9.65% of the time. The same documents under African-American-sounding names received a callback 6.45% of the time: roughly 50% more interview invitations for identical paper. Figure 2 shows the gap, and the gap is the point. Nothing a candidate wrote, earned, or did produced it.

The study contains a second finding that deserves equal weight. Improving résumé quality raised callbacks substantially for applicants with white-sounding names but much less for those with African-American-sounding names. Better credentials, the one signal a candidate can actually control, paid off unevenly along racial lines — the gap did not close with merit, it grew with it. For anyone designing a hiring process, that inverts the standard reassurance. A screening layer that rewards stronger qualifications only for some applicants is a filter in which merit and bias are entangled at the point of entry, not a meritocracy with noise sitting on top of it.

The result is not a single-study artifact, nor a peculiarity of two American cities. Philip Oreopoulos (2011) ran the same design at large scale in Canada, sending thousands of matched résumés to real vacancies, and found that applications carrying English-sounding names drew substantially more callbacks than identical documents carrying Indian or Chinese names. Two independent research teams, two countries, two sets of names, one pattern: the strongest signal traveling off the page is the one the applicant did nothing to earn.

These results describe racial discrimination, measured in the behavior of real employers, and they should be read at exactly that weight: neither inflated into a claim about every recruiter nor softened into a footnote about unconscious tendencies. The experimental design does not identify which screeners discriminated or why; it establishes that, in aggregate, the market did. The structural lesson is what a hiring leader can act on. Résumé bias operates at the stage with the most decisions and the least oversight, where rejected applicants vanish without feedback and no audit trail exists. An unvalidated judgment, exercised at volume, under time pressure, on a document whose strongest verified signal is demographic: nothing in that sentence describes measurement.

Biographical data can predict, when nobody is reading it

Here the story takes a turn that separates it from a simple indictment of biography. In the 2022 validity table, empirically keyed biodata stands at .38, among the stronger single predictors in the literature (Sackett et al., 2022). Biodata is biographical information: education, work history, experiences, activities, the same raw material a résumé contains. Everything hangs on that qualifier. Empirical keying means each item is scored by its statistically established relationship to later job outcomes, with the scoring key built on data and checked on fresh samples, then applied identically to every applicant. No narrative, no impression, no reader.

The contrast with eyeball screening is the same contrast the interview literature has already made famous within the field: the structured interview at .42 against the unstructured interview at .19 in the same 2022 table. Same conversation, opposite discipline, and more than half the validity gone; our companion article on why unstructured interviews fail traces that collapse in detail. Biography behaves identically. Scored against outcomes, it predicts at .38. Read impressionistically, it decays into the near-floor validities of its most legible lines, filtered through whatever the reader happens to weight: institutional prestige, an unbroken career story, formatting fluency, a familiar-sounding name. The information is not worthless. The eyeballing is.

This finding closes off the most tempting response to everything above, which is to train screeners to read résumés better. Reading is not the valid version of the activity; scoring is. And a genuine empirical key is a psychometric artifact, built from outcome data across many hires, cross-validated, and maintained, which is why almost no employer screening résumés by hand possesses one for the role in question. An organization that cannot build validated keys for its biography still has a direct route to the .4 neighborhood: stop inferring ability from life history and measure it, with instruments whose construction is documented and whose scoring never met the candidate.

Automation made résumé screening faster, not truer

When application volume exploded, organizations did not revisit what the résumé could measure. They automated the reading. In high-volume funnels, applicant tracking systems (ATS) perform the first cut, matching documents against keywords, degree requirements, tenure thresholds, and knockout questions. This solved the volume problem, which was real. It did not touch the validity problem, because a filter can only be as predictive as the criteria it encodes, and the criteria are the same weak proxies the validity literature priced decades ago, with exact-phrase matching layered on top.

Employers themselves say as much. In Hidden Workers: Untapped Talent, a Harvard Business School and Accenture study, Fuller and colleagues surveyed employers about their own screening technology: 88% agreed that qualified, high-skill candidates are vetted out of the process by their automated screening systems because they do not match the exact criteria in the job description (Fuller et al., 2021). The figure is self-report, a survey of employers rather than an audit of their systems, and it should be read as such. That is precisely what makes it striking. Organizations rarely testify against their own tooling, and here the overwhelming majority acknowledge that the machine rejects people it should not.

The ATS screening problems employers report are usually framed as configuration errors: too many required keywords, over-strict parsing, rigid Boolean logic. The deeper problem is what the configuration expresses. Job descriptions accumulate requirements the way attics accumulate boxes, and each one becomes a knockout rule: a degree for a role generations of incumbents performed without one, an arbitrary years-of-experience threshold, a tool name that a synonym fails to match. Requirements compound conjunctively, so each defensible-sounding line multiplies the exclusion. The result is a self-inflicted supply constraint. Organizations describe shortages in the same roles whose funnels they have configured to discard, on unvalidated criteria, a large share of the people who could do the work.

One objection surfaces here reliably: candidates embellish, so surely the real problem is that the résumé cannot be trusted. Embellishment exists, and industry surveys report rates that vary too widely to cite responsibly, which is why no prevalence figure appears in this article. But the argument does not depend on lying in either direction. A perfectly honest résumé is still a weak instrument, because the weakness lives in what the document can carry and how it gets read, not in the sincerity of its author. Verification handles dishonesty, later and cheaply. Nothing handles invalidity except a different instrument.

Move measurement to the front of the funnel

The design conclusion follows from the evidence with unusual directness. If the top of the funnel runs on near-noise plus bias, the remedy is not to refine the reading but to replace the layer: give every applicant a short, valid assessment before any human forms an impression of them, and let the résumé re-enter later in the role it can actually perform. Figure 3 contrasts the two architectures. In a résumé-first funnel, valid measurement is applied only to the survivors of an unvalidated cut, so whatever the assessment battery achieves downstream, the population it sees has already been filtered on the wrong variables. In a measurement-first funnel, the ordering reverses: the valid instrument makes the first cut, and human attention, the scarcest resource in hiring, is spent on candidates selected for measured signal rather than for document polish.

The standard objection is cost: assessing a thousand applicants sounds expensive next to skimming a thousand documents. It was, when every test was a fixed hour-long form. That constraint no longer binds. Because an adaptive test selects each next question for the information it carries about that specific candidate, it can match the precision of a much longer fixed form in a screening-length sitting; the mechanics, and what they imply for screening-length tests, are the subject of our companion article on adaptive testing in hiring. The candidate-burden objection inverts on inspection. An eyeball screen does not spare applicants effort; it denies most of them any hearing at all, and grants the hearing on the least valid and most bias-prone evidence available. A short assessment is the more respectful transaction: every applicant gets the same instrument, scored the same way, blind to name, template, and prestige.

What of academic proxies for candidates who have no work history to misread? Grades are the strongest case, and the evidence gives them a real but perishable signal. In a meta-analysis of grades and job performance, Roth and colleagues estimated a corrected correlation of about .3 shortly after graduation, weakening substantially as years since graduation accumulate (Roth, BeVier, Switzer, & Schippmann, 1996). That supports using GPA as one input among several in campus and graduate hiring, where it is recent and comparably normed, and it undercuts the degree-and-grades reflex everywhere else. An academic record from a decade ago is a proxy whose expiration date has passed.

The résumé, meanwhile, keeps the job it was always suited for. Degrees, dates, employers, and titles are verifiable facts, and checking them against records at the offer stage is cheap, accurate, and unhurried. The résumé is not evidence of quality; it is a set of claims — and claims are best checked late, when there are few candidates and real stakes, not mined for inferences early, when there are many candidates and no instrument in sight.

Treat the screening layer as the instrument it is

The screening layer produces scores. Advance or reject is a two-point scale, applied to every applicant, with consequences; it has reliability, validity, and adverse impact whether or not anyone measures them. The practical agenda is to hold it to the standard any other instrument in the process must already meet.

  1. Audit the screen against outcomes. An organization that would never buy an assessment without validity evidence runs résumé screening with none. Check what screening verdicts correlate with: later performance among those advanced, pass-through rates by demographic group, and the fate of borderline candidates. If the screen cannot be shown to predict anything, it is consuming candidates without producing information.
  2. Drop proxy requirements you cannot defend. Every degree requirement and years-of-experience threshold is a claim that the proxy predicts performance in this role. The published values for those proxies are .10 and .18. Keep a requirement only where a genuine occupational reason (a license, a legal mandate, a demonstrable knowledge floor) survives scrutiny; otherwise it is a self-imposed supply constraint with adverse impact attached.
  3. Structure whatever human review remains. Where regulations or role realities require reading documents, import the discipline that rescued the interview: defined criteria fixed in advance, name and identifying details masked, ratings recorded independently against the criteria rather than as overall impressions. Structure will not turn biography into a strong predictor, but it narrows the channel through which bias and taste enter.
  4. Verify late; infer never. Move claims-checking to the offer stage, where it is thorough and cheap per candidate, and strip quality inference out of the top of the funnel entirely. The screening question is not "does this document impress" but "does this person show the measured attributes the role demands," and only an instrument answers that.

Figure 4 summarizes the reallocation. Everything on the résumé sorts into two piles: claims that can be verified against records, and inferences a reader projects onto the page. The first pile is genuinely useful, at the right stage. The second pile is where the near-floor validities and the callback gaps live, and no reading technique converts it into measurement. The organizational move is not to read the second pile more carefully. It is to stop sourcing those judgments from the document and source them from an instrument built for the purpose, with construction and scoring a review panel can inspect; what such a screening layer looks like in operation is laid out in the platform overview, and what its output looks like per candidate in a sample report.

Biography is not meaningless, and the evidence never claims it is. What the evidence endorses is one particular version of it: scored, keyed, and verified, never skimmed. The organizations that internalize the distinction will not read résumés better than their competitors. They will simply stop asking a claims document to do an instrument's job, and their funnels will select on signal from the first applicant to the last.

Where 5Profiler stands

5Profiler is built to run screening in the order the evidence recommends: every applicant measured first, the résumé consulted afterward as the claims document it is. Adaptive testing is what makes that affordable at the top of a high-volume funnel, since each candidate's test adjusts to their answers and reaches decision-grade precision in a screening-length sitting. Nobody is cut by a template, a missed keyword, or a name before they have been measured. The first shortlist a hiring team sees is therefore selected on evidence it can inspect, and the verification work waits until the field is small enough for it to be done properly.

Read the science behind the platform · See it on your roles

References

  1. Bertrand, M., & Mullainathan, S. (2004). Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. American Economic Review, 94(4), 991–1013.
  2. Fuller, J. B., Raman, M., et al. (2021). Hidden workers: Untapped talent. Harvard Business School Project on Managing the Future of Work / Accenture.
  3. Oreopoulos, P. (2011). Why do skilled immigrants struggle in the labor market? A field experiment with thirteen thousand resumes. American Economic Journal: Economic Policy, 3(4), 148–171.
  4. Roth, P. L., BeVier, C. A., Switzer, F. S., & Schippmann, J. S. (1996). Meta-analyzing the relationship between grades and job performance. Journal of Applied Psychology, 81(5), 548–556.
  5. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
  6. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.