Research Scores & decisions

Will it work here?

The intuitive answer — run your own validity study — is a noise generator below several hundred hires. Validity generalization, transportability, synthetic validity and Bayesian updating: the evidence playbook for real org sizes.

Every validity number a vendor can show you was estimated on someone else's workforce. Assessment buyers notice, and the noticing is healthy: the coefficients in the research summaries come from other companies, other jobs, sometimes other decades, and the question they provoke is exactly the right one. Will it hold here? The trap sits inside the answer that feels most rigorous. Run a local validation study, the reasoning goes: test your own applicants, wait, correlate scores with the performance of the people you actually hired, and replace borrowed evidence with owned evidence. For most organizations, the number that comes out of that exercise is not owned evidence. It is sampling noise with a decimal point.

The problem is sampling error: a correlation computed on a small sample swings widely around the true value, and the swing does not announce itself. A validity coefficient estimated on a few dozen hires arrives with two decimal places and the full apparatus of measurement whether it landed near the truth or far from it. The field quantified this trap in 1976 and has been rediscovering it ever since, most often in the person of an employer who commissions a study, reads a null result, and concludes that a well-evidenced instrument fails on its people.

The escape runs through the evidence structure the profession has codified: validity generalization, which turns cumulative meta-analytic evidence into the starting estimate; transportability, which argues through structured job analysis that your job belongs to the family the evidence already covers; synthetic validity, which assembles evidence for a job out of evidence about its components; and Bayesian designs that let whatever local data exists update the cumulative estimate instead of pretending to replace it. Local evidence keeps real work in that structure. What it loses is top billing.

A local validation study at feasible size mostly measures luck

Statistical power is the probability that a study of a given size will detect an effect that is genuinely there, and power is the property the intuitive local study lacks. Schmidt, Hunter, and Urry (1976), writing in the Journal of Applied Psychology, computed the power of the standard criterion-related validity study under the conditions that actually obtain in employment data. Two of those conditions do most of the damage. Range restriction is the first: an employer observes job performance only for the people it hired, a narrowed slice of the applicant range, and correlations computed on narrowed slices come out smaller than the relationship that produced them. Criterion unreliability is the second: supervisor ratings, the default criterion, measure true performance with error, which shrinks the observed correlation again. Under realistic levels of both, the sample sizes required for adequate power sit far beyond what most single employers can assemble; for many realistic conditions the requirement runs to several hundred scored hires in a single role. The validation samples the field had historically been using were, by these calculations, badly underpowered.

The steepness of the requirement is visible even before those two spoilers enter. Figure 1 plots the number of cases a study needs for 80% power, the conventional benchmark meaning an 80% chance of detecting the effect, at a two-tailed significance threshold of .05, computed from the standard formula with no allowance for restriction or unreliability at all. To detect an observed validity of .20, an entirely plausible value once restriction and a noisy criterion have done their shrinking, a study needs about 194 cases. Even at an observed .40 the requirement is about 47, and 47 scored hires in one role, rated on one consistent standard, is more than many organizations can produce. These are the textbook numbers; Schmidt, Hunter, and Urry's point was that realistic conditions push the true requirement higher still, because the conditions shrink the very correlation the study is trying to detect.

Set that requirement against what a single employer holds. The cases that count are people who took the instrument, were hired anyway, stayed long enough to be rated, and were rated on a consistent standard. Assembling a file of that kind in the hundreds takes most organizations years, and the years are not neutral: the role drifts, managers turn over, the applicant pool shifts, and the accumulating sample slowly comes to describe a job that no longer quite exists. Meanwhile the study's audience rarely knows any of this. A committee shown a correlation from its own data will weight it above a meta-analysis of other people's, on the reasonable-sounding ground that local facts beat imported ones. The figure says otherwise: in the shaded zone, a local study is not measuring local facts at all. Samples that size produce coefficients dominated by chance, and reading one with conviction changes nothing about what produced it.

The upshot cuts both ways. A null result from an underpowered local study is, with high probability, a power failure, not a validity failure: the study could not have detected the validity even if the validity is real. And a small study that returns a gratifying coefficient has established just as little, because it is a draw from the same wide distribution, taken from the flattering tail. Both misreadings operate in practice. Instruments that work get dropped on null local results, and instruments that do not work get kept on lucky ones, and the organization cannot tell which of the two it has just done.

Why the field stopped demanding local proof

For decades the field's own doctrine sided with the intuitive buyer. Observed validities for the same test varied so much from study to study that the profession concluded validity was situationally specific (a test that worked in one plant might genuinely fail in the next one), and the prescription followed: validate locally, everywhere, every time. Schmidt and Hunter (1977), in the Journal of Applied Psychology, proposed what they called a general solution to the problem of validity generalization, and it dissolved the doctrine's foundation. The variation that had impressed everyone was substantially statistical artifact: sampling error above all, plus study-to-study differences in criterion unreliability and range restriction. Account for the artifacts, and much of the apparent situational variation resolves into noise around a far more stable underlying validity. Validity generalizes across settings to a degree the specificity doctrine never allowed.

There is a genuine irony in the sequence. The scatter that made local proof seem necessary was largely manufactured by the size of the local studies themselves: small samples produced wild coefficients, the wildness was read as situational specificity, and the specificity was cited to demand more small samples. Validity generalization broke the loop by pooling. Aggregate the studies, correct for the artifacts, and the cumulative estimate becomes steadier than anything any single employer could compute alone.

What validity generalization licenses is a change of posture. The cumulative meta-analytic estimate stops being vendor advertising to be discounted and becomes the prior: the estimate an informed adopter holds before any local data arrives. The current benchmark values for the major selection methods are the 2022 re-estimates (Sackett et al., 2022); the series briefing on the general mental ability debate examines that re-estimation in full, and the flagship ranking of what predicts job performance applies it method by method. The size of the prior is worth caring about because small differences in validity compound across every offer an organization extends, a translation the briefing on the return on investment of valid selection performs in currency.

What the posture leaves open is the question the next section answers. A cumulative estimate describes a family of jobs, and an employer holding the prior still owes an account of membership: why this role, in this organization, belongs to the family the meta-analyses pooled. Validity generalization relocates the burden of proof rather than abolishing it. The demonstration that remains is far smaller than a criterion study, and it is analytical work a modest HR function can actually complete, which is what makes the codified alternatives usable at every scale the local study fails at.

Transportability and synthetic validity carry the evidence to your job

Neither move described in this section is an improvisation; both are named strategies in the profession's governing documents. The Standards for educational and psychological testing hold that validity evidence may be assembled from multiple sources, local evidence being one source among several (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014). The SIOP Principles, the field's application of those standards to selection, list the accepted strategies by name: criterion-related studies, validity generalization, transportability, and synthetic or job-component approaches, with the expectation that the strategy chosen should match what is feasible as well as what is rigorous (Society for Industrial and Organizational Psychology, 2018).

Transportability is the argument that your job already belongs to the evidence. It runs through structured job analysis: document the tasks, the knowledge and skill demands, and the working context of the local role, and show that they match the jobs from which the cumulative estimate was built. When the match holds, the estimate travels, and the burden of proof sits where an employer can actually carry it, in job documentation. The match is claimed at the level of work, and it can fail; a shared job title establishes nothing about shared demands, and checking the demands is what the analysis is for. Done well, the exercise leaves a by-product more durable than its conclusion: a documented account of the role's actual requirements, which the organization will reuse every time the evidence question is asked again.

Synthetic validity assembles the estimate instead of transporting it whole. Decompose the job into components (numerical work, persuasion, vigilance, coordination), attach to each component the validity evidence that exists for it, and recombine according to the job's actual mixture. Johnson, Steel, Scherbaum, Hoffman, Jeanneret, and Foster (2010), writing in Industrial and Organizational Psychology under a title that states its verdict ("Validation is like motor oil: Synthetic is better"), made the case that evidence assembled this way can outperform a conventional local study. Components recur across jobs, so evidence about components accumulates at a scale no single employer's sample can approach, and a job that is novel as a whole is rarely novel in its parts. Hybrid roles, new roles, and roles that exist nowhere else are the natural clients, and they are exactly the roles a transportability match covers worst.

The routes differ most in what they demand, and Figure 2 compares them on that axis alongside the conditions under which each wins. The route that feels most like ownership carries the harshest entry conditions in the figure: cases in the hundreds, a criterion worth correlating with, and years of accrual before the coefficient stabilizes. Transportability and synthetic validity substitute analytical work for sample size, and the substitution is what makes them feasible at any scale. The bottom row changes how every other row is used, and it gets the next section to itself.

Bayes-analysis makes a thin local file count

The three-way contest has been run directly. Newman, Jacobs, and Bartram (2007), in the Journal of Applied Psychology, compared the accuracy of the local validity estimates an organization would reach by each available route: meta-analysis alone, a local study alone, and Bayes-analysis, the combination in which the cumulative meta-analytic estimate serves as the prior and the local data updates it. The combination was typically the most accurate. The stand-alone local study became competitive only at large sample sizes, exactly the sizes Figure 1 says most employers never reach.

The mechanism is precision weighting. The prior arrives carrying the pooled weight of the cumulative literature; the local sample arrives carrying whatever situational signal it genuinely holds, wrapped in sampling error proportional to its thinness. Updating weighs the two by their precision, so a thin file moves the estimate slightly and an ample file moves it substantially. Each pure strategy's failure gets corrected in the process: meta-analysis alone ignores real local information, and a local study alone treats noise as revelation.

The geometry of the move is in Figure 3. The local study, at typical size, contributes the wide curve: an estimate that could land almost anywhere in the plausible range. The meta-analytic prior contributes the tight one. The posterior, the updated estimate, sits near the prior, shifted toward the local result only as far as the local data can justify, and tighter than either input, because agreement between independent sources is itself information. Nothing is discarded and nothing is crowned.

This posture also repairs the all-or-nothing habit that runs through assessment adoption. Organizations tend to treat local data as either decisive or worthless: the study is commissioned and believed, or never run at all. Bayes-analysis gives thin evidence a permanent, proportionate role. The file far too small for a criterion-related validity study on its own is exactly big enough to be an update. And because updates accumulate, the frame rewards patience: each cohort of captured criteria narrows the posterior a little further, so the organization's estimate improves on a schedule no single study could match, without a launch, a consultant, or a threshold moment at which the evidence is suddenly declared sufficient.

What local data is still for

The first local job is criterion quality. Every strategy in Figure 2 eventually touches locally measured performance: the Bayesian update needs a criterion to update against, monitoring needs outcomes to monitor, and supervisor ratings carry error of their own. Defining performance for a role, and measuring it consistently enough to correlate with anything, is work no meta-analysis can do on an employer's behalf. It is also cheap to start and impossible to backfill; a criterion not captured this year is a case lost to every future update. Criterion work is where local knowledge is genuinely superior, too: no meta-analysis knows what effective performance means on this team, in this market, under this operating model, and every downstream estimate inherits the care taken here. The companion briefing on measurement error deals with the uncertainty around an individual candidate's score; the uncertainty in this briefing lives one level up, around correlations, and better criteria shrink both.

Local data also carries the monitoring. Pass rates, completion rates, and score distributions by role are local facts, and drift in any of them is local news no cumulative estimate can deliver; subgroup patterns belong to the same review, under the frame the briefing on the four-fifths rule governs. Reading those distributions well requires knowing what a reported score is and is not saying, the subject of the companion briefing on what test scores mean.

Local data also sets the operating point. Where a cut score sits is a local decision even when validity is cumulative: labor market, hiring volume, and the consequences of each kind of error all move the line, and the briefing on setting cut scores carries that argument. Validity says the score deserves weight; local pass-rate data says where the weight starts to bind.

How much of this apparatus an organization needs tracks its volume, and the thresholds that follow are heuristics, offered as rough guidance and not as findings. Below roughly 300 scored hires in a role, run transportability and synthetic evidence and monitoring, and treat a stand-alone local validation study as the noise generator Figure 1 shows it to be. Between roughly 300 and 1,000, the local file starts to carry real weight in a Bayesian update, and formalizing that update is worth the effort. Above 1,000, volumes reached mainly in high-volume enterprise and campus hiring, a local criterion-related study becomes informative in its own right, and even then it enters the estimate as heavy local data in an update, since keeping the prior in the frame loses nothing.

Adoption is when the local evidence program begins

The productive request to a vendor is the cumulative evidence and the bridge to it: the meta-analytic record behind each instrument, and the job-analysis material that lets your organization argue the transportability case for its own roles. A vendor offering instead to prove itself with a quick study on a season's worth of your hires is offering the underpowered design this briefing began with, presented as diligence. The cumulative record is the stronger evidence, and asking for it is also the easier request; it already exists.

At adoption, write the transportability argument down: the job analysis performed, the match claimed, the evidence relied on. That documentation is what reviewers, successors, and skeptics will ask for when the instrument's presence is questioned years later, and the briefing on building a standards-based assessment program supplies the governance frame it files into. Criterion capture belongs on day one of the same program: performance defined, measured, and stored from the first cohort, on a platform that keeps scores and later outcomes side by side, so the local file accrues by default rather than by project. The update then belongs on a schedule, revisiting the estimate as the file grows, in the frame of Figure 3, so the organization's answer to the opening question improves with every cohort it hires.

The posture asks for the habit of mind that the briefing on mechanical versus holistic judgment documents at the level of individual decisions. The vivid local particular feels like better evidence than the dull cumulative table, and it almost never is: a manager's story about one candidate and a coefficient from one small study are both small samples carrying more confidence than information. Trusting the cumulative record over the local anecdote is a single act of method that pays twice, once in how candidates are judged and once in how instruments are.

"Will it work here?" turns out to have a better answer than the study intuition reaches for. The organization that documents its bridge, captures its criteria, and updates a cumulative prior on a schedule ends up holding something no bespoke study can produce: an estimate of local validity that starts from everything the field knows and moves, case by case, toward everything the organization learns. The question was never too ambitious for evidence. The intuitive method was simply too small for the question.

Where 5Profiler stands

5Profiler publishes the evidence base behind each instrument on the science page: the cumulative research each measure rests on, stated plainly enough to be checked against the published literature this briefing cites. That documentation is where a transportability argument starts, and role-referenced scoring is the capability the job-analysis bridge runs through. An organization adopting the platform inherits a prior, and a prior exists to be updated: the hiring outcomes the platform captures from the first cohort onward become the local data the estimate has been waiting for.

Read the science behind the platform · See it on your roles

References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
  2. Johnson, J. W., Steel, P., Scherbaum, C. A., Hoffman, C. C., Jeanneret, P. R., & Foster, J. (2010). Validation is like motor oil: Synthetic is better. Industrial and Organizational Psychology, 3(3), 305–328.
  3. Newman, D. A., Jacobs, R. R., & Bartram, D. (2007). Choosing the best method for local validity estimation: Relative accuracy of meta-analysis versus a local study versus Bayes-analysis. Journal of Applied Psychology, 92(5), 1394–1413.
  4. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
  5. Schmidt, F. L., & Hunter, J. E. (1977). Development of a general solution to the problem of validity generalization. Journal of Applied Psychology, 62(5), 529–540.
  6. Schmidt, F. L., Hunter, J. E., & Urry, V. W. (1976). Statistical power in criterion-related validation studies. Journal of Applied Psychology, 61(4), 473–485.
  7. Society for Industrial and Organizational Psychology. (2018). Principles for the validation and use of personnel selection procedures (5th ed.). Industrial and Organizational Psychology, 11(Suppl. 1), 1–97.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.