A situational judgment test hands the candidate a workplace scene, a short list of plausible responses, and one question about them. Every part of the instrument can stay fixed (the scenarios, the options, the scoring key) while a single word in that question is swapped out. Ask what a person should do, and the test behaves like a measure of knowledge and cognitive ability. Ask what the person would do, and the same items behave like a personality inventory.
The swap has one strange property: it changes everything about the test except how well it works. In the meta-analysis that mapped this fork, both versions predicted job performance identically, to the second decimal place, observed and corrected alike (McDaniel, Hartman, Whetzel, & Grubb, 2007). One construct replaces another, and prediction never registers the difference. An instrument that predicts equally well no matter what it measures is not a marvel of robustness. It is a format with no fixed subject matter.
The SJT, as a class, is a method wearing a construct's name tag. What any given test measures is settled by design choices (the response instructions, the intended construct, the scoring key, the form in circulation this season) that most organizations never inspect. Made deliberately, those choices produce a useful, defensible instrument. Left to default, they produce a score that measures whatever the format happened to pick up.
Born as a simulation, not as a construct
The format entered the research literature as an engineering compromise. Motowidlo, Dunnette, and Carter (1990) wanted the realism of a work sample without the burden of building and staffing one, so they proposed what they named the low-fidelity simulation: written scenarios drawn from real incidents on the job, each followed by a set of response options, scored against a key derived from experienced managers' judgment. In their initial studies of entry-level managers, scores correlated roughly .28 to .37 with supervisors' ratings of later performance, a respectable showing for a paper exercise requiring a fraction of a simulation's apparatus.
The name conceded the trade openly. A low-fidelity simulation sits on the same continuum as the work sample, at the lighter end: the candidate judges a situation on paper rather than performing in one, exchanging behavioral realism for scale. What the high-fidelity end of that continuum makes possible is the subject of the companion article on work samples.
What the name did not do was specify a construct. The anatomy of an SJT is spare: a scenario, some options, a key. Typical situational judgment test examples in hiring stay close to the daily texture of work: an escalating customer, a colleague sliding on deadlines, two urgent requests and one afternoon. Nothing in that anatomy names the psychological attribute being scored. The attribute gets attached afterward, by whoever writes the items and whoever writes the brochure, and that sequence (format first, construct later) is the source of nearly everything odd in the evidence that follows.
Compare the build order everywhere else in assessment. A personality inventory begins from a trait taxonomy and writes items to cover it; a knowledge test begins from a syllabus; a cognitive battery begins from a definition of the ability it samples. Construct first, format second. The SJT reversed the order and prospered anyway, because the format is likable: candidates read scenarios that resemble the job, hiring teams see the relevance at a glance, and the whole thing ships as easily as a questionnaire. Popularity of that kind is evidence about face validity, the sense that a test looks right, and about nothing else.
What situational judgment tests predict, and where the signal comes from
Begin with the headline validity, then its footnote. McDaniel, Morgeson, Finnegan, Campion, and Braverman (2001), in a synthesis published to clarify a scattered literature, pooled 102 validity coefficients covering 10,640 people and estimated the SJT's corrected validity against job performance at .34. That number anchored the format's reputation for two decades. The current best estimate is lower: re-analyzing the selection literature with less aggressive range-restriction corrections, Sackett, Zhang, Berry, and Lievens (2022) put SJT operational validity at .26, and the story of why the field's estimates moved belongs to the companion article on the GMA debate. For scale, structured interviews stand at .42 and general mental ability tests at .31 in that same revision, with the full ranking charted in the series flagship.
A validity of .26 justifies a place in a selection battery but cannot carry one. It sits below the strongest single methods and comfortably above the résumé proxies most screening still runs on: enough signal to sharpen a battery's composite, not enough to select on alone. The more consequential number in the 2001 synthesis is a different one entirely: across 79 correlations and 16,984 people, SJT scores correlated .46 with general cognitive ability (McDaniel et al., 2001).
That overlap is not incidental. Reading a scenario, holding four courses of action in mind, and reasoning through their consequences is cognitive work before it is anything else, so a portion of any SJT's signal rides on the ability variance its scores carry. Whether that portion is a feature or a contaminant depends entirely on what the test claims to measure, and the claim is exactly what most SJTs leave vague.
The overlap also complicates the case for adding one. An SJT usually joins a battery that already contains an ability measure, and a supplement helps in proportion to the new information it brings. A .46 correlation means part of what a generic SJT contributes is ability signal the battery already owns, so its practical value hangs on the residual: the construct-specific variance left over once ability is accounted for. That residual is the part the designer controls, and the rest of this article is about controlling it.
The instructions decide what the format measures
In 2007, McDaniel, Hartman, Whetzel, and Grubb split the SJT literature along a line most test users have never been told exists: the response instructions. Knowledge instructions ask the candidate to identify the best possible response, what one should do; they elicit maximal performance, a display of what the person knows about effective behavior. Behavioral-tendency instructions ask what the candidate would most likely do; they elicit a self-report of typical behavior, which is the same kind of claim a personality inventory collects. Their meta-analysis covered 96 studies of knowledge-instructed SJTs (22,050 people) and 22 studies of tendency-instructed ones (2,706 people).
The two families turn out to be different instruments built from the same parts. Knowledge-instructed SJTs correlate .35 with cognitive ability; tendency-instructed versions correlate .19. The personality correlations invert: emotional stability correlates .35 with tendency-instructed scores against .12 with knowledge-instructed ones, and Figure 1 traces agreeableness and conscientiousness crossing over in the same way. The instruction is not an administrative detail. It selects which psychological attribute the scenarios will absorb.
Then the criterion refuses to care: knowledge-instructed and tendency-instructed SJTs showed the same mean observed validity of .20 and the same corrected validity of .26 (McDaniel et al., 2007). Two constructs, one criterion. Knowing what works and typically doing what works both predict performance, by different psychological routes, to the same degree. The tendency side of the literature is far smaller than the knowledge side, which counsels some caution, but the symmetry has stood.
Figure 1 doubles as a warning label. A vendor's claim that its scenarios "measure judgment" is compatible with either branch of the fork, and the branch is chosen by wording the vendor may treat as boilerplate. Before any other question about an SJT, ask which instruction family it uses, because the answer determines what construct the scores can legitimately be interpreted as.
The fork also constrains what a score report may legitimately say. A knowledge-instructed SJT supports statements about what a candidate understands about effective behavior in the scenario domain; it does not license inferences about temperament, however strongly the scenario content suggests them. A tendency-instructed SJT supports the opposite family of statements, at least outside high-stakes use. A report that translates one instruction family into the other's language, describing a should-test as a window into disposition or a would-test as demonstrated know-how, is making claims the instrument cannot back.
High stakes collapse the fork
The fork has a further twist, documented by a randomized experiment run inside a live, high-stakes selection program. Belgium admits students to medical school through a national exam whose battery includes an SJT, and Lievens, Sackett, and Buyse (2009) arranged for the response instructions to be randomly assigned: 2,184 applicants saw identical scenarios, some asked what they should do, the rest asked what they would do, with admission genuinely riding on the answers.
Under those conditions the instruction effect nearly vanished. The mean score difference between the two versions was d = .10, a tenth of a standard deviation. The correlation with cognitive ability, .35 versus .19 across the low-stakes literature, compressed to .19 versus .11. Criterion-related validity did not differ. Applicants facing real consequences answered would as if it read should: offered the chance to describe their typical behavior, they described their best behavior, which is what the knowledge version asks for outright.
The result slots into the broader evidence on self-report under stakes: an instruction that invites impression management will receive it whenever admission or employment hangs on the score. For SJT design the implication is blunt. The tendency instruction's personality-like character is partly a low-stakes phenomenon, so an organization cannot count on getting a personality measure merely by asking a would-question of applicants. The Belgian program drew the operational conclusion and adopted knowledge instructions permanently, reasoning that if candidates will answer should regardless, the test ought to be scored and defended as what it actually is.
The boundary of the finding matters in the other direction too. The collapse was produced by admission-level stakes. In lower-stakes settings (development programs, internal mobility diagnostics, practice assessments), the fork in Figure 1 presumably reopens, and a tendency-instructed SJT can recover its personality-like character. The same item bank can therefore behave differently at different points in the employee lifecycle, one more way the format's meaning travels with its context of use rather than with its scenarios. A program that runs one SJT for hiring and reuses it later for development should expect the two administrations to yield scores with different meanings.
A third of published tests cannot say what they measure
If instructions decide what an SJT measures, someone still has to decide what it is supposed to measure. Christian, Edwards, and Bradley (2010) audited the published SJT literature and classified every test by its intended construct. Their census, charted in Figure 2, found leadership the most common target.
The bar that matters most is the one without a name. The second-largest share of the published literature could be classified only as heterogeneous composites: tests reporting a single overall score that spans several constructs at once, with no interpretable construct attached. A composite of that kind cannot be normed against an external benchmark, compared across instruments, or defended when a rejected candidate asks what, exactly, the assessment measured. There is no construct claim to defend, and every downstream artifact of the score (norms, candidate feedback, appeal decisions) inherits the vagueness.
Construct choice is not cosmetic; it moves validity. The strongest construct-specific estimate in the meta-analysis belongs to teamwork-targeted SJTs at .38, though that figure rests on only six studies and should be read as promising rather than settled; estimates for the other construct families cluster lower, as annotated in Figure 2. Medium matters as well: video-based SJTs outpredicted paper versions of the same constructs (interpersonal skills .47 versus .27; composites .36 versus .25), with the video estimates likewise drawn from small study counts (Christian, Edwards, & Bradley, 2010). Both patterns lean the same direction. Validity sharpens when a test is built as a deliberate, behaviorally rendered measure of one construct, and blurs when it is built as a generic judgment test.
What a second attempt buys, and what it does not
A construct-first SJT still has to survive operational life: applicants who take it twice, coaching markets that grow around any consequential exam, and the awkward look of its own reliability statistics. The retest evidence again comes from the Belgian admission program. Lievens, Buyse, and Sackett (2005) tracked 1,985 people who took the full battery a second time and measured what the repeat attempt was worth on each instrument.
Figure 3 shows the outcome. Retaking raised SJT scores by d = .32 (.40 after correcting for predictor unreliability), a real gain, but a smaller one than the d = .42 the cognitive ability test showed in the same battery; Figure 3 sets both beside the knowledge test. The SJT is not unusually retest-sensitive among the instruments it typically accompanies. More important, SJT validity held: the test predicted subsequent performance equally well whether first-attempt or retest scores were used. Score inflation did not become prediction decay.
Retest gains of that size explain why coaching markets form around operational SJTs, and they point at the defense. A gain available to repeat takers is available, in some measure, to coached ones, and the countermeasures are the same for both: alternate forms, so familiarity with any single form loses its value; key security, so the only material left to coach is the construct itself; and validity checked across attempts, which the Belgian data show is feasible inside a live program. What the retest evidence does not support is the reflexive conclusion that a retaken SJT is a compromised one. The retest scores went on predicting.
The same research program supplies the reliability context that organizations routinely misread. Judged by internal consistency, SJTs look broken: coefficients around .30 to .40 across alternate forms, far below any conventional threshold. But internal consistency assumes a test's items all measure one thing, and SJTs are multidimensional by design; a single scenario can tax judgment, interpersonal reading, and self-control at once. The appropriate yardsticks are alternate-form and test–retest standards, and there the picture is merely modest: test–retest reliability of .66 in a one-week study of 250 takers (Lievens, Buyse, & Sackett, 2005). An SJT with a low alpha is not necessarily malfunctioning. An SJT program that reuses one form across seasons, with gains of a third of a standard deviation available for a second sitting, is.
Name the construct before you pick the format
For organizations, the practical sequence inverts the usual one: the construct comes before the vendor. Decide whether the role's evidence gap is interpersonal skill, leadership judgment, team coordination, or applied knowledge, and set aside any candidate instrument whose documentation cannot state what its score means. A test that answers the construct question with a vague appeal to judgment belongs to the unlabeled third of the literature, whatever its production values.
Instructions follow from the construct. Where the target is knowledge-like (procedural savvy, applied expertise, knowing what effective behavior looks like), knowledge instructions measure it directly and carry the larger evidence base. Where the target is genuinely dispositional, a tendency-instructed SJT is one option, but the Belgian experiment says candidates under real stakes will drift toward should no matter what the prompt asks, so an organization that wants personality measured should usually measure personality with a personality instrument and let the situational judgment test do a different job in the battery. How a multi-instrument battery divides that labor at hiring volume is described in the platform overview.
Two further disciplines come straight from the evidence. Hold the scoring key to the standard that anchored rating scales set for structured interviews: a documented, expert-derived rationale for why the keyed response is keyed, retrievable when a decision is challenged. And plan for reuse before the first administration, rotating alternate forms across hiring seasons, because retest gains are real even where validity holds. Fairness belongs on the same checklist, and here the format has a genuine attraction: demographic score gaps on SJTs are meta-analytically much smaller than on cognitive ability tests, and the gap grows with the SJT's cognitive loading, knowledge-instructed formats being the more g-saturated (Whetzel, McDaniel, & Nguyen, 2008), a trade-off that belongs in any adverse-impact review.
Documentation is where these requirements become checkable. Before deployment, an SJT's paper trail should state its intended construct, its instruction family, the provenance of its scoring key, and the plan for rotating forms, in language a non-psychologist on a hiring committee can act on. Each of those lines answers a failure documented above: the unlabeled composite, the fork, the indefensible key, the reused form. An instrument that arrives with all four stated has at least made its choices deliberately. An instrument that arrives without them has made the same choices by accident, and its scores will mean whatever the accidents add up to.
The Belgian admission program shows what the whole discipline looks like assembled. Its SJT ran inside a battery, beside a cognitive test and a knowledge test, never in place of either. When a design question arose, the program ran a randomized experiment on its own exam; when the experiment answered, it committed, adopting knowledge instructions permanently, maintaining alternate forms, and checking validity across repeat attempts. The scenarios never changed. What changed is that the program could say precisely what its judgment test measured, and could show it still measured the same thing on a candidate's second try.
Where 5Profiler stands
Construct-first is how 5Profiler builds judgment measurement. A scenario-based judgment module starts from a named construct and the role's actual incidents, with response instructions and scoring key chosen to fit that construct, so the score a hiring team receives says what it measures. Each module's report states that construct, its instruction family, and the provenance of its key, keeping the score's meaning on the page. And judgment modules run alongside purpose-built instruments, never in their place: where the question is personality, the platform measures personality directly, at 30-facet resolution, instead of asking scenarios to double as an inventory. The scenario stays a scenario, and the construct keeps its name.
References
- Christian, M. S., Edwards, B. D., & Bradley, J. C. (2010). Situational judgment tests: Constructs assessed and a meta-analysis of their criterion-related validities. Personnel Psychology, 63(1), 83–117.
- Lievens, F., Buyse, T., & Sackett, P. R. (2005). Retest effects in operational selection settings: Development and test of a framework. Personnel Psychology, 58(4), 981–1007.
- Lievens, F., Sackett, P. R., & Buyse, T. (2009). The effects of response instructions on situational judgment test performance and validity in a high-stakes context. Journal of Applied Psychology, 94(4), 1095–1101.
- McDaniel, M. A., Hartman, N. S., Whetzel, D. L., & Grubb, W. L., III. (2007). Situational judgment tests, response instructions, and validity: A meta-analysis. Personnel Psychology, 60(1), 63–91.
- McDaniel, M. A., Morgeson, F. P., Finnegan, E. B., Campion, M. A., & Braverman, E. P. (2001). Use of situational judgment tests to predict job performance: A clarification of the literature. Journal of Applied Psychology, 86(4), 730–740.
- Motowidlo, S. J., Dunnette, M. D., & Carter, G. W. (1990). An alternative selection procedure: The low-fidelity simulation. Journal of Applied Psychology, 75(6), 640–647.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- Whetzel, D. L., McDaniel, M. A., & Nguyen, N. T. (2008). Subgroup differences in situational judgment test performance: A meta-analysis. Human Performance, 21(3), 291–309.