Research Applied playbooks

Structure is a craft.

The evidence for structured interviews is settled; the practice is not. The field manual: job-derived questions, behaviorally anchored scales, independent ratings and rule-based scoring.

Derive the questions from the job's actual demands, put the same questions to every candidate, score the answers against behaviorally anchored rating scales, rate independently, and combine the ratings by rule, and a conversation becomes a measurement instrument. The evidence for the structured interview is settled; the practice is not. What most organizations run under the name is a question list, and a question list is the cheapest tenth of the method. The nine tenths that carry the validity are the parts hiring teams drop first.

The gap matters more than it looks, because the halfway version is not half as good. A question list without anchors, without independent ratings, and without a combination rule keeps the psychometrics of an improvised conversation while conferring the confidence of a validated method: unstructured validity, structured credit. Panels in this configuration believe they are running the method the evidence endorses. The evidence endorsed something else.

Why improvised interviews fail, and how far the free-form estimate has fallen, is settled territory, mapped in the companion article on unstructured interview validity. This article takes the question that verdict leaves open: structure wins, so what does running it involve? The answer sits at the working level most writing on hiring skips, and it begins where every structured interview begins, with the job.

What a structured interview technically requires

The technical meaning of "structured" is broader than the word suggests. Campion, Palmer, and Campion (1997), whose review in Personnel Psychology serves as the catalogue of what structure means in the selection interview, sorted the elements into two families: content elements that fix what is asked, and evaluation elements that fix how answers become scores. The catalogue runs long (limits on prompting, consistent panels, note-taking discipline), but five elements form the load-bearing chain. Questions are derived from an analysis of the job's actual demands. The same questions are put to every candidate. Answers are scored against anchored rating scales whose points are pinned to described behavior. Ratings are made independently, before any discussion. And the ratings are combined by a rule fixed in advance (Campion, Palmer, & Campion, 1997).

The stakes are on the record. When Sackett, Zhang, Berry, and Lievens (2022) revisited the meta-analytic evidence on selection methods, applying less aggressive corrections for range restriction than earlier syntheses had, structured interviews came out at .42, a validity coefficient, the correlation between interview scores and later job performance (Sackett et al., 2022). The chain itself is design logic rather than a description of any single study's procedure: each link closes a channel that improvisation leaves open, and the sections that follow build the links in order.

Figure 1 lays the chain out, with what each link is there to stop. Two links repay a closer look, because they are the ones panels drop most readily in practice: anchored scoring makes two raters mean the same thing by a 4, and independent ratings keep the first confident voice from setting everyone's memory of the answer. Break any link and the damage is not local: an interview with superb questions and adjective scoring simply delivers its noise later, at the rating step, and an interview with anchored scores and a free-form debrief delivers it at the final step instead.

Partial adoption is the configuration to fear, because it fails while reassuring. The typical real-world pattern keeps the first two links and drops the last three: a shared question list, improvised probing, unanchored ratings, and a debrief that folds impressions into a verdict. Everything the chain was built to stop re-enters after the questions are asked. What remains is an improvised interview with better paperwork, wearing the validity reputation of the disciplined one, and the reputation is what makes it dangerous: a validated label on an unvalidated practice ends the search for the problem.

Write the questions from the job, not from folklore

Every link downstream depends on the first. A question is job-derived when it traces to a demand the role actually makes, established by looking at the work: the recurring situations that separate a good quarter from a bad one, the incidents where the last person in the seat earned or lost their keep. It is folklore when it traces to an interviewer's biography: the greatest-weakness opener, the favorite brainteaser, the culture question that changes with the asker's mood. The working test is short. For each question, name the job demand it probes and the answer behavior that would count as evidence. A question that cannot pass is not a bad question; it is a question about a different job, usually an imaginary one.

The research offers two working formats. Latham, Saari, Pursell, and Campion (1980) built the situational interview around a specific discipline: pose a dilemma drawn from the job's real forks, ask what the candidate would do, and score the answer against anchored responses prepared in advance. The premise is that intentions, stated under the mild pressure of a concrete choice, carry signal about how a person will weigh the role's competing goods. The dilemma must genuinely fork: a situational question with an obvious right answer is a courtesy, not a measurement.

The second format asks for the past instead of the hypothetical. Experience-based questions, studied by Pulakos and Schmitt (1995), put the same job demands to the candidate as history, and in their studies the experience-based format performed at least as well as the situational one. The two share the properties that matter, job-derived content and anchored scoring (Pulakos & Schmitt, 1995), so the choice between them is practical rather than doctrinal. Experience-based questions need a past to interrogate, which mid-career candidates have and fresh graduates largely do not; in early-career pipelines, the kind examined in the companion article on campus hiring at scale, the situational format does the interviewing that experience cannot yet support.

Figure 2 builds both formats from one illustrative demand, ours, for a coordination-heavy role: keeping commitments when priorities collide. The situational version constructs the fork and asks the candidate to choose under it; the experience-based version asks for the last time the fork actually arrived. Note what each reveals. The situational answer exposes reasoning: which value the candidate protects, what they say to the parties involved, whether the plan survives the constraint. The experience answer exposes conduct: what was actually protected, what actually slipped, and, in the best probes, what the miss cost and who heard about it first. A strong loop uses both formats on the same demands and scores both against anchors.

Follow-ups are where structure usually dies in the room. The fix is not to ban probing, which would discard evidence, but to plan it: each question ships with its intended probes (what did you do next, what did it cost, who else was involved), and the panel stays inside them. Planned probes are part of the question. Improvised ones are a path back to a different interview per candidate, and the difference between the two is invisible in the transcript unless the probes were written first.

An anchor describes behavior a rater could have heard

Scoring is where the remaining structure most often fails, and the failure has a standard look: a five-point scale labeled poor to excellent, attached to a question that took a week to write. The labels feel like rigor because they are numbered. They are, in fact, the absence of a scale. "Good" does not tell two raters whether the same answer earns a 3 or a 4; it tells them to consult the impression the whole apparatus was built to bypass, and the impression obliges.

The repair is older than most of the interview literature. Smith and Kendall (1963) framed the approach that became behaviorally anchored rating scales: instead of numbering a feeling, pin each scale point to a description of behavior, then keep only the descriptions that independent judges reliably re-sort to the level they were written for, a procedure they called retranslation (Smith & Kendall, 1963). Retranslation is the quality-control step modern rubric writing usually skips, and it is cheap to approximate. Draft the anchors, hand them unsorted to colleagues who know the role, and keep the ones everyone files at the same level. An anchor filed at 3 by one reader and at 5 by another is not an anchor yet. It is a future argument between raters, caught early enough to fix.

Three craft rules do most of the discriminating. Anchors describe observable behavior at each level, in terms a rater could match to what was actually said: what a strong answer contains, what a middling one contains, what a weak one omits. They cover the real range, including the middle, because a scale with vivid endpoints and an empty center pushes every uncertain rater to the same safe 3. And they never lean on evaluative synonyms, because "excellent" is a verdict, and the anchor's job is to make the verdict computable from evidence rather than the reverse.

Figure 3 shows the difference on the same illustrative competency the questions were built from, with anchor text that is ours and drawn from no study. The left column is the anti-pattern: five labels restating the judgment the scale was meant to discipline. The right column gives described behavior at three levels of a five-point scale, which in practice is enough: anchor the points you can describe unambiguously, and let raters place an answer between described levels rather than stretching a vague description across every point. This is what practitioners mean by a BARS interview, and the anchor set doubles as the interview scoring guide: the one document a new panelist can be trained on in an afternoon, because it says what to listen for instead of how to feel.

The library compounds. Anchors are written per competency, and competencies recur across roles: the anchor set for keeping commitments serves the coordinator, the project lead, and the account manager with modest edits, which is why the second role is cheaper to structure than the first and the fifth is nearly free. Teams that treat anchor writing as a per-requisition cost overestimate it; teams that treat it as a library underestimate what they already own after a quarter of doing it.

The panel rates alone before it talks

Everything the chain has bought is still forfeitable at the debrief, which is why the evaluation elements in Campion and colleagues' catalogue end with two disciplines that belong together: ratings made independently, and ratings combined by rule. The first governs when opinions may meet. The second governs what happens after they have.

Independence first. Each panelist scores every competency against the anchors before any conversation, in writing, with the supporting evidence noted. Once ratings exist on paper, discussion can do what discussion is good for: surface an answer one rater caught and another missed, flag a question that misfired, correct a fact one panelist misheard. What discussion can no longer do is set the group's memory of the answers before anyone has committed to a judgment, which is the same channel through which bias enters a debrief, the case the companion article on how structured hiring reduces bias makes in full.

Combination last, and by rule. Kuncel, Klieger, Connelly, and Ones (2013) meta-analyzed what happens when the same information, identical scores, ratings, and records, is combined mechanically, by an explicit rule, rather than folded into a holistic overall judgment: combining by rule improved predictive validity by more than half (Kuncel et al., 2013). For a panel the instruction is concrete. Weights across competencies are fixed when the loop is designed, the composite is computed from the written ratings, and discussion informs a documented reading of the evidence rather than a renegotiation of the score. Re-rating stays possible: a panelist who learns she misheard an answer changes her rating, with the reason logged. What the protocol removes is the unlogged version, where scores drift toward the room's center of gravity and no one can later say why.

Two smaller elements from the same catalogue earn their keep here. Keeping the panel consistent across candidates means ratings do not drift with the cast, and taking notes during the interview, rather than reconstructing answers from memory at the debrief, gives the independent ratings something firmer than recollection to rest on (Campion, Palmer, & Campion, 1997). Both cost minutes. Both are usually the first casualties of a crowded interview week.

Figure 4 puts the sequence on a timeline, and the line that matters is the first one: no scores are shared until every rating is locked. Discussion then interprets the record rather than overwriting it, the composite is computed from the written ratings, and the whole file, ratings, notes, and any logged re-rates, is kept. The protocol reads as ceremony until the first time a panel runs it, at which point something surprising surfaces almost immediately: the raters disagree more than anyone expected, and the disagreement is information. A debrief-first panel would have converged on the senior rating and called it consensus. A rate-first panel gets to ask which anchor the disagreeing raters were looking at, which is a question with an answer.

After an assessment, the interview stops re-testing the file

A structured interview is expensive per candidate-hour, which makes the question of what the hour is for a budgeting question and not a philosophical one. Loops without upstream evidence use that hour to re-derive what a test measures better: intelligence inferred from articulateness, diligence inferred from a firm handshake and a prompt arrival, both inferences colored by whatever the file primed. The interview can be a fine instrument and still be pointed at the wrong target.

An assessment battery upstream changes the target. When cognitive ability and the candidate's personality profile arrive measured, under standardized conditions a conversation cannot reproduce, the interview is released from re-measuring them and can be aimed where measurement runs out. The report tells the panel where that is, and it does so candidate by candidate. A profile with an unusually strong result invites verification in behavior: an experience-based question aimed at where that strength should have left footprints. An ambiguous or borderline result invites concentration: the interview aims its minutes exactly where the file is least certain. And the demands no test reaches well, domain judgment, stakeholder craft, the texture of past work, get the remaining time as the interview's primary and proper subject. The probes attached to a candidate report show the shape this takes in practice: the assessment turns the interview from a re-test of the file into a probe of what the file cannot show.

The sequencing matters as much as the content. A panel that reads the report after forming its impressions uses the data to arbitrate a fight the impressions already started; a panel that builds its probes from the report before the loop begins never starts the fight. This is the same design logic that runs through the whole chain, applied at the level of the funnel: decide what the conversation is for before the conversation happens, and let each instrument, test and interview alike, do the work the other cannot. How the interview slots into the rest of an assessment-led funnel is laid out on the platform overview.

Audit the chain before you buy anything new

The practical program follows the chain in order, and it starts with an audit rather than a purchase. Most organizations do not need a new interview method; they need to find out which links of the one they claim to run actually exist.

  • Audit every loop against the five links. For each role, ask where the questions come from, whether every candidate gets the same ones, what raters score against, when ratings happen relative to discussion, and how ratings become the decision. Most loops pass the second link and fail the last three. Knowing which links are missing converts a vague ambition to interview better into a work order.
  • Write anchors for the three roles you fill most often. Anchor writing is the costliest link, so spend the effort where the repetitions are. Three high-volume roles anchor most of the interviewing an organization actually does, and because competencies recur, the library built for them seeds every role that follows.
  • Train raters on the anchors, not on technique. Rapport and question delivery improve interviews at the margin; calibration improves them at the core. Have new panelists rate the same sample answers against the anchor set and argue the discrepancies down before they meet a live candidate. When two raters file the same answer two levels apart, the anchor, or the rater, gets fixed before it costs a decision.
  • Keep the ratings. Retained, structured ratings are what make a loop improvable and defensible: they can be checked against outcomes later, audited for drift, and produced when a decision is challenged. An interview that leaves no record was, for institutional purposes, an impression.

The program asks for no new headcount and no new vendor; it asks the organization to treat the interview as what the evidence says it can be. The price is paid mostly in discipline rather than money, and what it buys is rare: the one step in hiring that is also a human conversation, keeping its warmth and earning its decision weight at the same time.

Where 5Profiler stands

A 5Profiler report reaches the panel with the interview already half-built. Attached to each candidate's profile are interview probes aimed at what the assessment could not settle, so the conversation verifies and extends the evidence instead of re-testing it across a table. Around the probes sits the discipline this article describes: behaviorally anchored rubric libraries that give raters described behavior to score against, and panel scorecards that take each rater's judgment independently, then combine the scores by rule once every rating is in. Probes, rubrics, and scorecards are all in place before the candidate walks in, which is where this article has argued an interview is actually built.

Read the science behind the platform · See it on your roles

References

  1. Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3), 655–702.
  2. Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98(6), 1060–1072.
  3. Latham, G. P., Saari, L. M., Pursell, E. D., & Campion, M. A. (1980). The situational interview. Journal of Applied Psychology, 65(4), 422–427.
  4. Pulakos, E. D., & Schmitt, N. (1995). Experience-based and situational interview questions: Studies of validity. Personnel Psychology, 48(2), 289–308.
  5. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
  6. Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2), 149–155.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.