Research Cost of getting it wrong

The interview illusion.

Free-form interviews predict performance at .19 — and in controlled studies they made predictions worse than no interview at all. Why the most trusted method in hiring earns the least trust, and what fixes it.

"The scores are strong, but I just didn't feel it." One sentence, offered late in a hiring debrief by the most senior person at the table, and the decision turns on it. The candidate cleared a cognitive test, performed well on a work sample, and matched the role's trait profile; none of that survives the remark. The research on unstructured interview validity says the room has this exactly backwards. The impression doing the overruling is the weakest signal in the file, and in controlled studies it is worse than weak: it degrades the predictions of the very people who hold it.

The unstructured interview survives on a paradox. The people most confident in their ability to read a candidate across a table are the people the evidence says should trust that ability least, and decades of hiring interview research have documented the mismatch from every angle. Free-form interviews sit low in the validity rankings. Interviewers who watch the same candidate reach different verdicts. The questions asked bend toward impressions formed before the candidate sits down. And adding an interview to a file of standardized evidence reliably raises decision-makers' confidence without raising their accuracy.

The interview itself comes through intact. The same body of research shows that structuring the conversation more than doubles its predictive validity and lifts it to the strongest single predictor in the most recent meta-analytic synthesis. The argument of this article is therefore not "stop interviewing." It is "stop improvising": stop treating a free-form conversation as a measurement instrument, and stop letting the impressions it produces veto better data collected earlier in the funnel.

Unstructured interview validity has fallen with every re-analysis

Selection research prices every hiring method in a single currency: the validity coefficient, a correlation between scores at hire and job performance measured later. Our flagship article on what actually predicts job performance walks through the full rankings and the machinery behind them. Here, only two rows of the table matter, and both of them describe an interview.

Schmidt and Hunter's 1998 synthesis in Psychological Bulletin, which compressed decades of selection studies into one comparative table, put the unstructured interview at .38 and the structured interview at .51 (Schmidt & Hunter, 1998). Read charitably, .38 was respectable; plenty of methods scored lower. But the gap was the story. The same conversation, run with discipline, predicted substantially better, which meant the difference was never the medium. It was the improvisation. Interviewers were right that conversations carry information. They were wrong about how much of it survives an unstructured format.

Then the estimates were revisited. Sackett, Zhang, Berry, and Lievens (2022), writing in the Journal of Applied Psychology, re-examined the statistical corrections beneath the classic numbers, specifically the adjustment for range restriction (validity studies can only observe people who were actually hired, a score-compressed group, so observed correlations are corrected upward). Applying less aggressive range-restriction corrections, they cut the unstructured interview's estimate roughly in half, to .19, while the structured interview eased to .42 and moved to the top of the revised table (Sackett et al., 2022). What the revision did to every other method is the flagship article's subject; what it did to the free-form interview is this one's.

Figure 1 lines up structured against unstructured interviews across both syntheses, and the pattern repeats in each generation of evidence: whatever else the two research eras disagree about, structure carries the signal. Nothing about the room, the participants, or the length of the conversation changes between those two rows. Only the discipline does.

An unstructured interview validity of .19 does not mean interviews never catch anything; a weak signal is still a signal. But weak signals earn their place only when they arrive cheap and do not displace stronger ones, and the free-form interview fails both tests. It consumes senior staff hours one candidate at a time, and, as the next section shows, its output does not politely take its place alongside better evidence. It dilutes it. That combination is the modern funnel's strangest feature: the step placed nearest the decision, and trusted most, is its least reliable instrument, and it is routinely permitted to overrule the strongest.

A conversation can subtract: the dilution experiment

The sharpest evidence against the free-form interview is not that it adds little. It is that it can subtract. Dana, Dawes, and Peterson (2013), in a set of studies published in Judgment and Decision Making, asked participants to predict something unambiguous and checkable: students' actual semester GPA. Some forecasters received biographical information about the student. Others received the same biographical information plus an unstructured interview with the student, the design a hiring panel would assume gives the richer picture.

The forecasters given the interview predicted less accurately than the forecasters given the biographical record alone (Dana, Dawes, & Peterson, 2013). The conversation did not merely fail to add signal; it diluted the valid information the forecasters already held. Vivid impressions from the encounter were blended into the judgment, taking weight that belonged on the duller, more diagnostic record. And then the detail that explains every debrief you have ever sat in: the participants nonetheless preferred having the interview. The forecasters liked the condition that made them worse, which is why Dana and colleagues titled the phenomenon "the persistence of an illusion."

Figure 2 shows the design and the direction of the result, and direction is all the figure claims; the sign, not the size, is the scandal. Translate the design into hiring terms and the debrief from the opening reads differently. The interviewer who "just didn't feel it" is not contributing an independent channel of evidence to sit beside the test scores and the work sample. On this evidence, the feeling is substantially noise, and blending it into the decision waters down information the organization already paid to collect. A companion article prices what a diluted hiring decision costs when it goes wrong; the point here is where the dilution enters.

Why does a conversation dilute rather than enrich? The mechanics are mundane. An interview generates material that is vivid, recent, and personal, and human judgment gives vividness more weight than it gives diagnostic value. A flat line in a biographical record and a warm exchange across a table do not feel like comparable evidence even when the record is the better predictor, so the judgment tilts toward the encounter. Nothing about expertise switches this off, which is the finding the next section makes uncomfortable in detail.

Why the noise feels like signal

For the dilution result to change practice, it has to survive a natural objection: surely that is other interviewers. Skilled ones, the objection goes, extract real information from a free-form conversation. The difficulty is that the collapse in unstructured interview validity is not a mystery about skill. Hiring interview research has identified the machinery, and the machinery operates on everyone.

Start with agreement. If a free-form verdict were mostly signal, two interviewers watching the same candidate should land in roughly the same place. Conway, Jako, and Goodman (1995) meta-analyzed interrater reliability in selection interviews (the degree to which different interviewers score the same candidate the same way) and found that agreement rises substantially with structure, roughly doubling from low-structure to high-structure formats. Read in reverse, that is a measurement of idiosyncrasy: strip the structure away and a large share of the verdict reflects who happened to conduct the interview rather than who sat across from them. A measurement that changes with the measurer is not reading the candidate. It is reading the interviewer.

Second, the evidence an interview collects is not independent of the interviewer's prior. In a field study of real employment interviews, Dougherty, Turban, and Callender (1994) found that interviewers' pre-interview impressions of a candidate shaped how they conducted the meeting itself. Candidates the interviewer already favored received more selling of the company and the job, less questioning, and more confirmation-seeking; candidates viewed skeptically got the interrogative version of the same hour. This is interview bias operating upstream of judgment. The free-form interview does not gather evidence and then form an impression; it forms an impression and then manufactures the evidence to fit.

Third, the experience of interviewing feels like learning even when it is not. Kausel, Culbertson, and Madrid (2016) gave decision-makers standardized predictor information about candidates, then added unstructured interview impressions to the file. Confidence rose more than accuracy, and the added certainty was miscalibrated (Kausel, Culbertson, & Madrid, 2016). Figure 3 sketches the pattern conceptually: adding the conversation to a standardized file feels like resolution, but the confidence line climbs while the accuracy line stops following it.

These three gears mesh into a self-sealing system. Low agreement means the verdict is personal; confirmation-seeking means the conversation flatters the prior; miscalibrated confidence means the interviewer walks out more certain and no more correct. And because rejected candidates disappear, the loop never closes: an interviewer's gut is almost never scored against outcomes, so it cannot lose. Highhouse (2008) argued that this is precisely why intuition persists in selection despite the evidence — practitioners hold a durable belief in their own interviewing skill, and nothing in ordinary hiring practice ever confronts that belief with a scoreboard. The illusion is not a personal failing. It is the default output of a process that supplies vivid experience and withholds feedback.

What the interview is actually for

It would be easy to read this literature as a case for abolishing interviews, and that reading would be wrong. The strongest case for the interview deserves stating in full. Hiring is mutual: the candidate is deciding too, and the interview is where an organization makes its case, answers hard questions, and treats a person with the seriousness the decision deserves.

A good interview also delivers a realistic preview of the role, the team, and the manager, which candidates are entitled to before uprooting a career. It surfaces logistics, expectations, and mutual dealbreakers that no instrument measures. And it is the moment of human contact in a process that would otherwise feel like being processed; removing it signals indifference to the people an employer most wants to persuade. The ritual carries organizational value beyond prediction, too: a team that met its new colleague before the offer went out has a stake in making the hire succeed.

The research is also narrower than its sharpest summaries suggest. Dana and colleagues' forecasters predicted GPA in a deliberately clean laboratory task, not a full hiring funnel; field interviews serve several purposes at once, and an experienced interviewer holds genuine knowledge about the job and the team. The evidence does not say conversation contains no information, and it does not say interviewers know nothing.

It says something more precise: free-form impressions are unreliable across raters, contaminated by priors, and overweighted by the people who form them, so they should not function as a measurement instrument, and they should never silently outvote instruments that are one. The structured interview exists exactly because the conversation contains signal worth extracting; at .42 it outpaces every other single method in the 2022 estimates. The waste is not in talking to candidates. It is in improvising at them.

The fix is boring, and that is the point

What separates a .19 conversation from a .42 conversation is documented in unusual detail. Campion, Palmer, and Campion (1997), reviewing the accumulated research on interview structure in Personnel Psychology, catalogued the elements that carry the improvement. Four do most of the work: questions derived from an analysis of the job's actual demands; the same questions put to every candidate; answers scored against anchored rating scales, meaning scales whose points are pinned to described behaviors rather than adjectives; and ratings made independently, before any group discussion (Campion, Palmer, & Campion, 1997). Figure 4 shows the sequence.

Each element closes one of the failure channels from the previous section. Job-derived questions are how an interviewer's real knowledge of the role enters the process legitimately, as content rather than as hunch. Identical questions deny the Dougherty mechanism its lever, since a favored candidate can no longer be handed an easier meeting than a doubted one. Anchored scales pin scores to what was actually said instead of how it felt. Independent ratings protect the room from its loudest prior, which is the debrief failure the opening described. Run this way, the conversation stops measuring the interviewer and starts measuring the candidate.

The standard objection is that structure makes interviews robotic. It constrains the wrong thing for that worry to hold: the discipline lives in what is asked and how answers are scored, not in how warmly the meeting is conducted. A structured interview can open with rapport, sell the role, and leave generous time for the candidate's own questions; what it cannot do is let the small talk grade the candidate. Candidates, for their part, are asked the questions the job actually turns on, and are scored on their answers rather than on an interviewer's mood, which is a fair trade from both chairs.

Structure has one more step to run, because the gains of disciplined measurement can be forfeited at the last moment, in how the pieces are combined. Kuncel, Klieger, Connelly, and Ones (2013) meta-analyzed what happens when selection and admissions decisions combine the same information (identical scores, ratings, and records) either mechanically, by explicit rule, or holistically, by judgment. Combining by rule improved predictive validity by more than half (Kuncel et al., 2013). The implication for the debrief is blunt: a panel that runs a structured interview and then folds the ratings into an open-ended conversation about overall feel has rebuilt the unstructured interview at the final step and given back much of what the structure bought.

A structured debrief is therefore as specifiable as a structured interview. Before anyone speaks, each interviewer submits ratings on the anchored scales, competency by competency. The ratings are combined by a rule fixed in advance, alongside the standardized evidence: test scores, work samples, job-relevant trait profiles.

Discussion then audits evidence rather than renegotiating verdicts: what was asked, what was answered, whether a score has support. Disagreement is still welcome; it simply has to arrive as evidence about answers rather than adjectives about people. Overrides of the composite remain possible, but they are written down, reasoned, and tracked against outcomes, so the organization gradually learns what its exceptions are worth. The only thing that disappears is the sentence this article opened with, because a feeling with no evidence attached stops counting as an argument.

Run interviews late, structured, and without a veto

For organizations willing to act on the evidence, the operating rules are few and concrete.

  1. Sequence the funnel by validity, not by tradition. Standardized measurement (ability, work samples, personality profiles matched to the role's demands) should filter the pool before anyone books a meeting; the interview then probes a short slate deeply instead of screening a long one shallowly. The résumé pass this replaces has accuracy problems of its own, examined in the companion piece on what résumé screening actually measures.
  2. Structure every interview, with no exemption for seniority. Job-derived questions, identical across candidates, anchored rating scales, and written independent ratings before any discussion (Campion, Palmer, & Campion, 1997). Seniority is where exemptions get requested, and it is where they cost the most, because the senior voice is the one a debrief defers to.
  3. Combine mechanically, and log every override. A rule fixed in advance makes the first decision (Kuncel et al., 2013); a free-form impression never outvotes the composite by default. Exceptions stay possible, documented, and audited, which converts overrides from vetoes into data.
  4. Train the discipline, not the mystique. Interviewer training should build skill at asking, probing, anchoring, and rating rather than promising sharper intuition. The belief in one's own candidate-reading ability is the root Highhouse (2008) identified, and structure is the working substitute for it.

The cost side of the ledger argues the same way from the other direction. A free-form veto that rejects a strong candidate wastes everything the funnel spent to find them; one that admits a weak candidate converts into salary, ramp time, management load, and an eventual regretted exit, arithmetic we price in the companion article on the cost of a bad hire. Interview hours are scarce, senior-priced time, and structure is how that time measures instead of confirms. Organizations that want to see what the standardized layers of such a funnel hand the panel can inspect a sample candidate report; how the assessment layer feeds a shortlist into a structured, independently rated debrief is laid out in the platform overview.

The evidence does not ask anyone to stop meeting candidates. It asks the meeting to earn its decision weight the way every other instrument in the funnel must: by measuring the same things, the same way, for everyone. The unstructured interview failed that test in the syntheses, in the laboratory, and in the field. The structured interview passes it, at .42. Between those two facts sits a choice every hiring organization makes, usually without noticing it has been made.

Where 5Profiler stands

5Profiler is built for the sequencing this evidence recommends: standardized measurement first, structured conversation last, and no free-form veto in between. Candidates complete the battery before the panel convenes, so every interviewer opens the debrief holding the same evidence on every candidate rather than a résumé impression and a memory of the room. Role-referenced scoring supplies that shared record. Ratings are then entered independently, on anchored scales, before discussion starts, and the composite is assembled by rule rather than renegotiated by whoever speaks with the most conviction. Overrides stay available to the hiring manager; they are recorded as decisions instead of absorbed as instinct. The interview stays in the process. The improvisation does not.

Read the science behind the platform · See it on your roles

References

  1. Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3), 655–702.
  2. Conway, J. M., Jako, R. A., & Goodman, D. F. (1995). A meta-analysis of interrater and internal consistency reliability of selection interviews. Journal of Applied Psychology, 80(5), 565–579.
  3. Dana, J., Dawes, R., & Peterson, N. (2013). Belief in the unstructured interview: The persistence of an illusion. Judgment and Decision Making, 8(5), 512–520.
  4. Dougherty, T. W., Turban, D. B., & Callender, J. C. (1994). Confirming first impressions in the employment interview: A field study of interviewer behavior. Journal of Applied Psychology, 79(5), 659–665.
  5. Highhouse, S. (2008). Stubborn reliance on intuition and subjectivity in employee selection. Industrial and Organizational Psychology, 1(3), 333–342.
  6. Kausel, E. E., Culbertson, S. S., & Madrid, H. P. (2016). Overconfidence in personnel selection: When and why unstructured interview information can hurt hiring decisions. Organizational Behavior and Human Decision Processes, 137, 27–44.
  7. Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98(6), 1060–1072.
  8. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
  9. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.