Research Economics of hiring

The money math of better hiring.

A formula published in 1949 converts assessment validity into currency — and it has survived every honest attack on its assumptions. The math, the caveats, and why selection quality is still systematically underpriced.

In 1949, the psychometrician Hubert Brogden published a short paper in Personnel Psychology with a title that reads like a promise — "When testing pays off" — and a single line of algebra that converts the accuracy of a hiring method into dollars. More than 75 years later, every serious estimate of the ROI of hiring assessments still runs through Brogden's equation. It has survived formalization, application at federal scale, and a generation of hostile auditing that cut its estimates by an order of magnitude. The one thing it has struggled to survive is being shown to managers, and that failure turns out to be the most instructive finding in the whole literature.

The equation's persistence is earned, because it asks for nothing exotic. It wants to know how many people you hire, how well your selection method predicts performance, how much performance varies among people doing the same job, how selective you can afford to be, and what the assessment costs. Every one of those quantities has since been measured, some of them across decades of validation research. The practice of expressing selection value in money, known in the field as utility analysis, is therefore not a rhetorical exercise. It is arithmetic on observed inputs.

The arithmetic is lopsided because the inputs are lopsided. Output varies enormously between workers holding the same job title; validity converts a defined share of that variation into value captured at the point of hire; and testing costs are trivial next to salaries. This article works through the model, its measured inputs, and a fully labeled example, then gives the two most damaging critiques their full weight, because a case this lopsided can afford them. The conclusion survives both. The scandal of utility analysis is not that its numbers are inflated; it is that even the deflated numbers are ignored.

One line of algebra prices the entire hiring funnel

Brogden's result, later formalized and extended by Cronbach and Gleser (1965) in their decision-theoretic treatment of psychological tests and personnel decisions, states that the annual gain from a selection procedure is ΔU = N × r × SDy × z̄ − C. The notation is compact; the logic is conversational. Figure 1 lays out the anatomy, and each term deserves one plain sentence.

N is the number of people you hire in a year. Selection value arrives one hire at a time, so whatever a better method is worth per hire gets multiplied by volume. r is the validity coefficient, the correlation between scores on the method and later job performance; the flagship briefing in this series treats it in full.

SDy is the term managers have never been asked to estimate, and it does the heaviest lifting: the dollar value of one standard deviation of job performance, which is roughly the gap between an average performer and a genuinely strong one. z̄ is the average standardized assessment score of the people you actually hire, and it is a function of selectivity: hire one applicant in ten and z̄ is high, hire seven in ten and it collapses toward zero. C is what you spend on assessment, and it is the only term on the cost side of the ledger.

Two properties of the model matter for everything that follows. First, it is linear in validity: each additional point of r buys the same dollar increment as the last, so improving an already decent method pays as well as rescuing a poor one. Second, cost is subtracted, not multiplied. A method that costs more but predicts meaningfully better does not dilute the return; it shaves the C term while growing the entire gain side. Selection procedures should therefore be compared on net utility, not on price, a distinction most procurement processes invert.

Notice, too, what ΔU actually measures: a comparison, not an endorsement. The model prices the gap between a proposed way of choosing among the same applicants and the incumbent way, whether the incumbent is a weaker method or selection that is effectively random. That framing is what makes the equation useful in a budget meeting. The question is never whether assessment is worth something in the abstract; it is whether the specific gain in validity on offer, at your hiring volume and your selectivity, clears the specific cost being quoted.

None of the inputs is hand-waved

Take the terms in order, r first. Meta-analysis, the pooling of many validation studies into one estimate, is how the field prices r, and the most consequential recent one is by Sackett, Zhang, Berry, and Lievens (2022), who re-examined the selection literature's headline coefficients in the Journal of Applied Psychology, applying less aggressive range-restriction corrections than their predecessors. Two of their estimates frame this article's example: structured interviews predict job performance at r = .42, while unstructured interviews sit at .19. The full ranking of methods is the subject of a companion briefing on what actually predicts job performance; the point here is narrower. Validity is an observed quantity with a documented range, not a vendor's adjective.

The distance between those two coefficients is the commercial fact hiding in the meta-analytic literature. Both numbers describe an interview, with the same candidates, comparable interviewer time, and the same intent; what separates them is process. Structure means consistent questions, anchored scoring, and evidence recorded against criteria rather than impressions, and it more than doubles what the conversation can see. Validity, in other words, responds to design. That is what makes it a purchasable input rather than a fixed property of testing, and it is why the r term belongs in procurement conversations at all.

SDy looked, for decades, like the model's weak joint: how do you put a price on a standard deviation of performance? Schmidt, Hunter, McKenzie, and Muldrow (1979) supplied the estimation rule that made utility analysis practical, in a Journal of Applied Psychology study that has anchored the literature since: the standard deviation of the dollar value of output can be taken as about 40% of annual salary. Their demonstration case, computer programmers in the U.S. federal workforce, implied productivity gains approaching $100 million per year, in 1979 dollars, under top-down hiring assumptions. The figure was startling, and it was meant to be.

The 40% rule is not a universal constant, and its authors never claimed it was. Hunter, Schmidt, and Judiesch (1990) later showed that output variability scales with job complexity: the standard deviation of output runs about 19% of mean output in low-complexity jobs, about 32% in medium-complexity jobs, and about 48% in high-complexity and professional work (Figure 2). The managerial translation runs straight back into the model. SDy is not one number an organization looks up once; it scales with the work, so the same selection method applied to professional roles returns roughly two and a half times what it returns on routine ones. Complexity is a multiplier on every dollar in the equation.

This is the quiet engine of the whole model. Selection utility is large not because tests are magical but because people differ, and the more cognitively demanding the work, the more they differ. An organization that hires mostly for complex roles is running a high-variance process whether or not it manages that variance. Utility analysis merely writes the variance down.

A worked example, with every assumption in view

Utility analysis earns distrust when its assumptions hide, so here is a deliberately ordinary scenario with every input on the table. It is illustrative, not empirical. A company makes 200 hires a year at an average salary of $60,000. The 40% rule prices SDy at $24,000 per hire-year. The company attracts enough applicants to hire the top 30% of its pool, which under top-down selection yields a mean standardized score among those hired of z̄ ≈ 1.16. It currently screens with unstructured interviews (r = .19) and is weighing a structured process (r = .42), a validity gain of Δr = .23.

The Brogden arithmetic is one line: .23 × $24,000 × 1.16 ≈ $6,400 per hire, per year. Across the cohort, the switch is worth roughly $1.28 million per cohort-year before assessment costs. Note the unit, because it is easy to misread. That is the annual value of a single year's intake; the gain does not expire when the next requisition opens, it recurs for as long as those 200 people stay, and next year's cohort adds its own. Price C generously to close the ledger: assess 2,000 candidates, roughly three times the applicant flow the example needs, at $25 per candidate, and C comes to $50,000, about 4% of the gross gain, leaving a net near $1.23 million. The pricing is illustrative, like every other number in this scenario.

Every assumption here is contestable by design. Top-down selection assumes offers go out in score order and are accepted; real funnels leak. The 40% rule is a broad-brush convention rather than a bespoke costing of these particular roles. Salary is a convenient base, not a sacred one. But an exposed model converts each objection into a reprice instead of a rebuttal: cut SDy in half and the cohort gain halves with it, and the sign does not move. A reader who distrusts an input is invited to change it. The arithmetic, not the author, defends the conclusion.

Figure 3 generalizes the example across the two levers an organization actually controls: the validity of its method and the selectivity of its funnel. The lines are straight because the model is linear, and they fan out because selectivity multiplies validity. At r = .40, a top-10% hirer (z̄ ≈ 1.76) captures about $16.9K per hire-year; the same validity captures about $11.1K at top-30% selectivity and about $4.8K at top-70%. The marked point places the structured process on this map measured against no validity at all, not against the unstructured incumbent: r = .42 at top-30% selectivity captures about $11.7K per hire-year, and the worked example's $6,400 is the gap between that point and the unstructured interview's position on the same line.

Two readings follow. Selectivity is worth money, but only assessment makes selectivity affordable: distinguishing the top tenth of a large applicant pool by interviews alone is a scheduling problem measured in staff-weeks, while a well-built battery ranks the pool before anyone books a meeting. And validity pays at every level of selectivity, which means the r you buy, the measurement quality of the method itself, is a procurement decision with a computable return.

The ROI of hiring assessments survives its harshest critics

Numbers like $1.28 million invite suspicion, and the field, to its credit, supplied its own auditors. Boudreau (1983) argued that a serious economic estimate cannot stop at Brogden's gain term: future gains should be discounted to present value, taxed like any other profit, and netted against the variable costs that rise when output rises. Each correction is individually reasonable, and each shrinks the estimate.

Sturman (2000) then performed the demolition systematically. In a study published in the Journal of Management, he applied five families of adjustments to utility estimates, cumulatively rather than one polite correction at a time, and watched the numbers fall. The median estimate dropped by roughly 91%. An order-of-magnitude haircut amounts to a verdict that the headline numbers of classical utility analysis flatter themselves. But Sturman's analysis carries a second finding that is quoted far less often: after every adjustment, the estimates generally remained positive.

Run the worked example through that discipline and the $1.28 million becomes roughly $115,000 per cohort-year (Figure 4). The honest floor sits an order of magnitude below the textbook headline, and still comfortably above what a structured, well-instrumented selection process costs to run. A defensible ROI of hiring assessments does not require the optimistic estimate; the pessimistic one carries the decision on its own.

The second critique is more unsettling because it concerns persuasion, not arithmetic. Latham and Whyte (1994) — in a paper titled, without mercy, "The futility of utility analysis" — found that managers presented with a utility analysis were less committed to adopting the valid selection procedure than managers given validity information alone. The dollar figures did not close the sale; they undermined it. The plausible reading is that implausibly large numbers register as advocacy. A seven-figure return attributed to a line item most executives file under paperwork triggers precisely the skepticism it deserves.

Both critiques should be accepted rather than argued around. Sturman's is a correction to the estimate; Latham and Whyte's is a correction to the messenger. Together they define the right posture for anyone selling selection internally: compute conservatively, present modestly, and let the residual do the persuading.

Why better selection stays underpriced anyway

Return to the equation and look at C, because C is where organizational attention actually lives. Assessment spend is negotiated, benchmarked, and budget-lined; it has an owner, a renewal date, and a procurement file. Yet in the model, and in reality, C is the smallest object on the board. In the worked example, one standard deviation of performance is worth $24,000 per hire-year; against a number like that, per-candidate assessment pricing is not the decision variable. Negotiating C while ignoring r optimizes the term that barely matters. The ROI of hiring assessments is, at bottom, the ratio between the largest quantities in the equation and the smallest.

The deeper reason selection stays underpriced is that the variance itself appears nowhere in financial reporting. Payroll states the mean with precision: every salary, to the dollar. The spread around that mean, which Hunter, Schmidt, and Judiesch (1990) measured at 32% of mean output in medium-complexity jobs and 48% in high-complexity ones, has no line item, no owner, and no quarterly review. It is plausibly the largest untracked figure in the P&L. Utility analysis is best understood not as a sales tool for tests but as the accounting entry that variance never received.

A common reflex treats this as unactionable: people vary, and what of it? But the variance is addressable at exactly one moment, before the offer letter. After onboarding, closing a performance gap means training, managing, or replacing, each of which is slow, expensive, and uncertain. Selection is the only intervention that operates on the distribution itself, at the only point where the distribution is still a choice. That asymmetry is why the model concentrates so much value in a decision most organizations spend a few hours making.

Brogden's model is also conservative about the downside. It prices the average gain from better selection and says nothing about the tail, where the real cost of a bad hire accumulates: manager hours, team throughput, the attrition of adjacent performers, the eventual expense of re-hiring. Those mechanisms compound over tenure and across teams, and they deserve their own accounting, which the companion briefing on the cost of a bad hire supplies. For present purposes the point is directional. Everything the model leaves out makes the case stronger, not weaker.

Present the floor, and let the CFO ask about the ceiling

Latham and Whyte did not show that utility analysis is wrong; they showed that it is routinely mis-presented. The practical craft, for a talent leader walking into a finance review, is to present selection utility the way finance presents everything else: conservatively, in ranges, and anchored to numbers the audience already trusts. Four practices follow from the evidence.

  1. Lead with the floor, not the headline. Present the adjusted figure and name what was subtracted: discounting, taxes, variable costs, and the rest of the Sturman (2000) apparatus. The reverse order, as Latham and Whyte (1994) demonstrated, loses the room.
  2. Use ranges, not point estimates. The fan in Figure 3 is the true shape of the claim: value as a function of validity and selectivity, with the organization's own position marked. A range concedes uncertainty without conceding the sign.
  3. Anchor to metrics finance already tracks. Ramp time, first-year attrition, quality-of-hire reviews: better selection moves numbers the business already audits. Utility analysis should explain the mechanism behind those movements, not substitute for them.
  4. Treat validity as a purchasable input. r is not fate. Structure outpredicts improvisation (.42 against .19); multi-method batteries measure more of the candidate than any single instrument; and measurement precision itself is a specification you can demand: adaptive testing holds precision while shortening the sitting. Method quality belongs in the RFP, next to price.

For organizations hiring at scale, the fourth practice is where the model's linearity pays off: N multiplies every gain, so method quality compounds fastest exactly where enterprise hiring programs already apply procurement discipline. The standards worth demanding from any vendor are documented validity logic, transparent scoring, and evidence a psychometrician could audit. The demand matters more than the vendor. Assessment ROI is only as real as the measurement underneath it, and a utility model fed by a weak instrument prices a fiction.

Brogden's equation is old enough to have watched every generation of hiring technology claim to change the rules. Its verdict has been steady throughout: technology matters exactly insofar as it moves validity, reach, or cost, and validity was always the main event. The algebra was never the hard part. The hard part is the discipline of taking a conservative, adjusted, deliberately unexciting number seriously, and noticing that after decades of downward revision it has, in the main, stayed positive.

Where 5Profiler stands

The argument above is that validity is a purchasable input; the corollary is that it should be verified like one, on the buyer's own numbers. That is how 5Profiler is designed to be evaluated: pilot it in your own hiring, score it against outcomes you already track, and let the utility case be computed from your hires per year, your salaries, and your selectivity, not from a vendor's assumptions. Role-referenced scoring is the mechanism that keeps the r term pointed at the job you are actually filling. The arithmetic in this article is deliberately conservative. Filled with your figures instead of ours, it is the version of the case worth taking to a CFO.

Read the science behind the platform · See it on your roles

References

  1. Boudreau, J. W. (1983). Economic considerations in estimating the utility of human resource productivity improvement programs. Personnel Psychology, 36(3), 551–576.
  2. Brogden, H. E. (1949). When testing pays off. Personnel Psychology, 2(2), 171–183.
  3. Cronbach, L. J., & Gleser, G. C. (1965). Psychological tests and personnel decisions (2nd ed.). Urbana: University of Illinois Press.
  4. Hunter, J. E., Schmidt, F. L., & Judiesch, M. K. (1990). Individual differences in output variability as a function of job complexity. Journal of Applied Psychology, 75(1), 28–42.
  5. Latham, G. P., & Whyte, G. (1994). The futility of utility analysis. Personnel Psychology, 47(1), 31–46.
  6. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
  7. Schmidt, F. L., Hunter, J. E., McKenzie, R. C., & Muldrow, T. W. (1979). Impact of valid selection procedures on work-force productivity. Journal of Applied Psychology, 64(6), 609–626.
  8. Sturman, M. C. (2000). Implications of utility analysis adjustments for estimates of human resource intervention value. Journal of Management, 26(2), 281–299.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.