Research Ability & skills

Watch them work.

Work samples held the top of the validity table for a generation; the modern estimate is .33. And the method measures what a candidate can do today — not what they will do daily. The evidence, and the design choices that decide whether a simulation earns its cost.

Believing your own eyes feels less like judgment than like observation, and no evidence in hiring exploits that feeling like a work sample test: the panel watches the code run, the till balance, the demo land, and the debrief arrives pre-settled. Yet the try-out answers exactly one question — what this person can do today, at full effort, with skills they already hold. Hiring habitually reads further answers into the same performance: what the person will do daily once nobody is watching, and what they could learn if the job taught them. The evidence keeps those quantities apart, at distances that should change how simulations are used.

The distances are measured. On the field's revised numbers, work sample scores predict rated job performance at .33, upper tier among selection methods (Sackett, Zhang, Berry, & Lievens, 2022), so the can-do answer is real. The will-do answer is far looser: when researchers compared supermarket cashiers' do-your-best simulation scores with unobtrusive measures of their ordinary daily work, the correlations ran from .11 to .32 (Sackett, Zedeck, & Fogli, 1988). And the could-learn answer is not in the instrument at all. A work sample presupposes the skill it samples; it can only be given to people who already hold the job's abilities, and that restriction decides which applicant pools the method fits.

A work sample is a measure of maximum performance, what a person produces at their best under observation, while hiring keeps reading it as a measure of typical performance, what the same person produces on an ordinary unobserved Tuesday. Keeping those quantities separate decides how much fidelity a selection process actually needs, and it is the difference between hiring the demonstration and hiring the employee.

Why watching someone work predicts anything at all

The theory underneath work samples is older than most of the numbers in this series and plainer than any of them. Reviewing the early try-out literature, Asher and Sciarrino (1974) argued for a point-to-point account of prediction: a selection method forecasts job performance to the extent that its content shares points of correspondence with the job's content, the same behaviors, elicited by the same kinds of materials, in the same kind of context. Nothing about traits or aptitudes needs to be inferred. Resemblance itself carries the prediction, and the more points of contact between test and job, the higher the validity should climb.

Callinan and Robertson (2000) turned that account into the taxonomy the field still uses, organized by fidelity: the extent to which an assessment task and its context mirror the actual job. Every method sits somewhere on the resulting spectrum, which Figure 1 arranges. At the far left, a written knowledge test asks the candidate to describe the work. At the far right, probation stops simulating altogether; the assessment is the job, at full salary and full risk, with the verdict arriving months into employment. Everything hiring calls a simulation lives between those poles, and where it sits is a design decision, not an accident.

Reading the spectrum as a spectrum, not as a menu of interchangeable products, changes what the choice means. Each step rightward adds correspondence with the job, and the point-to-point account predicts that the added correspondence should show up as validity. Each step rightward also costs more to build, staff, administer, and score consistently, so a job simulation assessment is an operational decision before it is a psychometric one. The rails of Figure 1 rise together. The open question is whether they rise at the same rate.

The taxonomy also comes with boundary conditions, and Callinan and Robertson (2000) are careful about them. A sample stands in for the job only when it samples the job's discriminating demands, the tasks that separate good performers from poor ones. It stands in only when conditions are controlled enough that differences in scores reflect differences in people. And it stands in only when the behavior sampled is stable enough to be worth forecasting from. A simulation can miss all of these while looking magnificently realistic, which is the first hint that fidelity and value are different axes.

Validity survived the audit; the aura did not

For a generation, the number attached to that logic was .54. In Schmidt and Hunter's (1998) synthesis of 85 years of selection research, work sample tests stood as the top-ranked method on the whole table, ahead of cognitive ability and structured interviews, and the ranking fit the theory so well that it hardened into folklore: nothing should predict the job like the job. The figure's weakness was its foundation. It rested on a thin base of early studies, and revisiting that base was the stated motivation of the update that followed.

Roth, Bobko, and McFarland (2005) reassembled the work sample literature in Personnel Psychology: 54 studies covering 10,469 people. The mean observed correlation between work sample scores and job performance was .26. Corrected for one artifact only, the unreliability of the criterion (supervisor ratings and objective output are themselves noisy gauges of true performance), the estimate rose to .33. It stopped there deliberately. Fifty-three of the 54 studies were concurrent, meaning the test was validated on current employees rather than on applicants followed over time, so the authors applied no correction for range restriction at all: the setting gave no license for one. The design detail matters beyond statistics. A literature built almost entirely on current employees is a literature about people who already hold the skill, a point that returns below.

The broader re-analysis agreed. Sackett, Zhang, Berry, and Lievens (2022), re-estimating selection validities with less aggressive range-restriction corrections than earlier syntheses had applied, adopted the .33 outright as the field's operative estimate of work sample validity and endorsed the no-correction decision behind it. The fall from .54 to .33, a drop of .21, was among the largest revisions on their table. It leaves work samples below structured interviews at .42 and beside general mental ability at .31; the full reordering is traced in the series' review of what predicts job performance, and the argument over the corrections themselves runs through the general mental ability debate.

Read correctly, the audit is not a demotion. A .33 keeps the work sample in the upper tier of single methods, the point-to-point logic survives untouched, and no other format puts the criterion itself in front of the assessor. What the audit removed is the aura: the sense that watching someone work settles the hiring question, and everything else in the battery is garnish. At .33, the try-out explains a meaningful slice of performance and leaves most of it unexplained. The next study says what kind of slice it is.

The register data and the simulation disagree

Every coefficient above describes how well try-out scores predict rated performance. What no coefficient says is what kind of performance the try-out itself is, and the cleanest answer comes from a supermarket. Sackett, Zedeck, and Fogli (1988) studied cashiers in two samples of 635 and 735 people across 12 grocery chains, and measured each of them twice. Typical performance came from four weeks of unobtrusive data generated by the registers during ordinary work: speed, as items processed per minute, and accuracy, as voided transactions. Maximum performance came from a standardized checkout simulation each cashier completed knowing they were being evaluated, instructed to do their best.

If a work sample test were a window onto daily work, the two measurements would track each other closely. They track loosely at best. Among new hires, do-your-best speed correlated with day-to-day speed at .14; among current employees, at .32; across the four comparisons in Figure 2, the values run from .11 to .32. These are the same people, doing the same tasks, on the same machines. The gap between the numbers is not noise. It is the difference between what a person can do when it counts and what the same person does when it merely recurs.

The vocabulary the study fixed in the literature sorts this cleanly. Maximum performance is short-duration, observed, full-effort output: precisely the conditions of every audition ever staged. Typical performance is long-duration, unobserved output, and over unobserved weeks the binding constraint stops being capability and becomes motivation, the portion of capacity a person routinely applies. A work sample test is a maximum-performance instrument by construction. It reads the ceiling, and the ceiling is real information. The daily number is set somewhere below it, by dispositions the audition is structurally unable to see.

The split between new hires and current employees is itself informative. Among the new hires, whose skills were still forming, best-effort performance said almost nothing about daily output; among experienced cashiers the association strengthened but stayed modest. Even where skill has had years to consolidate, most of the variation in daily work is set by something the simulation never touches. And checkout work is a favorable case for measurement: short cycles, objective counts, little ambiguity about what speed and accuracy mean. The interpretation problem has no reason to shrink in jobs with longer arcs and murkier criteria; it only gets harder to see.

Cashiers are one setting, and the fair cross-study summary is friendlier to the audition without rescuing it. Beus and Whitman (2012) meta-analyzed the typical–maximum relationship and estimated it at ρ = .42, with the strength moderated by task complexity, the type of measure, and the setting. Related, then, and reliably so; interchangeable, never. Hiring on a work sample alone is a decision to weight what a candidate can produce under scrutiny and to guess at everything after onboarding.

Fidelity comes in increments, and the increments are small

The spectrum's remaining question is incremental: what each added layer of realism contributes that the layer below it does not already capture. The cleanest design in the literature comes from high-stakes medical selection. Lievens and Patterson (2011) followed 196 trainees selected into UK general practice, with supervisor-rated job performance measured about 12 months in and with three predictors scored on the same people against the same criterion: a written knowledge test, a low-fidelity situational judgment test built from text scenarios, and a high-fidelity simulation in which candidates performed the work live.

The zero-order validities, each corrected, land close together: knowledge test .54, low-fidelity SJT .56, high-fidelity simulation .50. Realism did not raise the headline number; the simplest instrument and the most elaborate one bracket the middle format within a few points. The study's information is in the increments, and Figure 3 stacks them. The knowledge test alone explained just under 30% of the variance in rated performance. Each simulation format, added to the knowledge test, contributed about 6% of variance on top; stacked after both the knowledge test and the SJT, the high-fidelity layer added 2–3% more.

The result cuts in both directions. Simulation is not theater: about 6% of variance is a real, non-redundant contribution, evidence that watching candidates handle work adds information a written test cannot reach. But the second layer of realism, the kind that needs actors, assessors, and facilities, contributed a fraction of what the first layer had already delivered. Fidelity obeys diminishing returns, and the diminishing starts early.

Increments are also relative to the battery, not absolute properties of a format. What each simulation layer added, it added on top of a strong knowledge test; against an emptier battery the same instrument would carry more, and stacked onto a fuller one, less. That is the correct frame for any claimed validity in this category: the question is never what a job simulation assessment explains on its own but what it explains beyond the instruments already in the battery. The GP study is valuable because it evaluates the layers on the same people against the same criterion, which is exactly the comparison that decision needs.

That increment-first reading is what the fidelity spectrum was missing. The low-fidelity middle, scenario-based judgment exercises deliverable at scale, captures most of what simulation as a category adds here; the evidence on situational judgment tests examines that format's record on its own terms. What a full assessment center adds, and the long-running puzzle of what it actually measures, is taken up in the review of the assessment center evidence. The general rule survives both discussions: add fidelity when an increment justifies it, and treat realism that cannot show one as production value.

A work sample test presupposes the skill it samples

Everything above carries a precondition so obvious it goes unexamined: the candidate must already be able to attempt the work. Schmidt and Hunter (1998) stated it as a scope condition, not a footnote: work samples can only be administered to applicants who already know the job. The same correspondence that generates the validity generates the restriction. An exercise built from the job's content is unreadable to a person the job has never trained, and a score of zero from that person is a fact about their history, not their capacity.

This is the line between a skills assessment test and an aptitude measure, and the 2022 re-analysis draws it as a matter of conceptual fit: ability tests suit untrained, entry-level hiring, where the question is who will acquire the skill, while work samples suit experienced hires, in whom the skill already exists to be sampled (Sackett et al., 2022). Point a work sample at a pool of fresh graduates and it returns a ranking of prior exposure, internships, hobby projects, and accidents of curriculum, printed with the false authority of a performance score. The error is invisible in the room, because the exercise still produces neat differentiation: some graduates do better than others. It differentiates on the wrong construct, and no amount of scoring rigor repairs a sampling decision.

The boundary shapes practice at both poles. The could-learn question needs instruments built for it, which is the architecture campus hiring programs run on. And where applicant pools reliably hold the skill, the work sample becomes the anchor instrument: the review of coding assessments works that application in full, including the standardization a coding task must meet before the general coefficients in this article can be claimed for it.

Decide which question you are asking before choosing the method

The failures documented here are failures of interpretation, so the guidance is mostly about matching, and it starts before any method is shortlisted. Name the question the role is actually posing. If the pool cannot yet do the work, the can-do question is empty, and the effort belongs in ability and learning measurement, not in simulation. If the pool can do the work, a work sample test answers the can-do question at .33 and answers nothing else. If the operational worry is consistency, effort, or conduct over months, that is a will-do question, and no amount of fidelity converts a maximum-performance instrument into an answer for it.

Then climb the fidelity spectrum the way the GP study reads it: from below. Start with the simplest format that can carry the construct, and require every step up in realism to defend itself with incremental validity, variance explained beyond the instruments already in place, not with face validity, the feeling of realism that impresses candidates and committees in equal measure. On the available evidence, a text scenario that captures the job's judgment demands is not a lesser version of the live simulation; it is most of the live simulation, at a fraction of the cost.

Realism does have one legitimate claim of its own. A lifelike exercise shows the candidate the job while the job screens the candidate, and organizations reasonably value that preview and the self-selection it produces. Adding fidelity for candidate experience is defensible; it is a recruiting decision, and it should be justified as recruiting, not claimed as measurement the increments cannot support.

Standardize as though the coefficient depends on it, because it does. The validities in this article describe fixed tasks, identical conditions, and anchored scoring; an improvised try-out, graded by whoever had the free hour, inherits the label and none of the evidence. The engineering of that standardization is worked through for technical hiring in the coding assessment review linked above, and its principles transfer to any domain, including platform-delivered exercises where identical administration is a property of the software rather than of the proctor's vigilance.

Finally, pair the audition with typical-behavior measurement instead of promoting it to sole witness. The cashier study supplies the reason: skill sets the ceiling, and the dispositions that govern where effort goes set the daily number. Measures of conscientiousness at work and related traits are typical-performance instruments in exactly the sense a work sample is not. That complementarity is the central idea behind multi-method assessment batteries: can-do and will-do measured separately instead of one being guessed from the other.

The case for work samples has always been that seeing is believing. The evidence agrees, more precisely than the slogan wants: what you see is what this candidate can do this week, at full attention, with the stakes visible. That is worth measuring, and at .33 it is measured well. But most of a job is the weeks after the watching stops, and those weeks are governed by what the person brings when no one is scoring. Hire on the audition alone and you have hired a ceiling — the daily life lived under it is what the rest of the battery is for.

Where 5Profiler stands

Skill demonstrations on 5Profiler run as standardized instruments: the same task, the same conditions, and the same time limits for every candidate, with scoring against anchored rubrics rather than a reviewer's impressions. That standardization is the condition under which the validity evidence in this article applies at all, and the platform enforces it in software, from delivery through scoring. The demonstration then enters the candidate's profile as one instrument among several: 30-facet personality resolution carries the dispositional measurement beside it, and both readings reach the hiring team in a single report.

Read the science behind the platform · See it on your roles

References

  1. Asher, J. J., & Sciarrino, J. A. (1974). Realistic work sample tests: A review. Personnel Psychology, 27(4), 519–533.
  2. Beus, J. M., & Whitman, D. S. (2012). The relationship between typical and maximum performance: A meta-analytic examination. Human Performance, 25(5), 355–376.
  3. Callinan, M., & Robertson, I. T. (2000). Work sample testing. International Journal of Selection and Assessment, 8(4), 248–260.
  4. Lievens, F., & Patterson, F. (2011). The validity and incremental validity of knowledge tests, low-fidelity simulations, and high-fidelity simulations for predicting job performance in advanced-level high-stakes selection. Journal of Applied Psychology, 96(5), 927–940.
  5. Roth, P. L., Bobko, P., & McFarland, L. A. (2005). A meta-analysis of work sample test validity: Updating and integrating some classic literature. Personnel Psychology, 58(4), 1009–1037.
  6. Sackett, P. R., Zedeck, S., & Fogli, L. (1988). Relations between measures of typical and maximum job performance. Journal of Applied Psychology, 73(3), 482–486.
  7. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
  8. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.