Research The AI era

The interview with nobody in the room.

Algorithmic interview scores lean on what candidates say, not how they look — and the medium itself lowers ratings (d = −.41). The evidence on automated video interviews, and the questions that separate instruments from demos.

The newest interview room holds one person. A candidate sits alone in front of a laptop, reads a question off the screen, and answers into the camera, with no interviewer to nod, redirect, or follow up; the listening happens later, when a model transcribes the recording and converts the words, the delivery, and sometimes the face into scores. AI video interviews did not simply move the conversation online. They changed what an interview is: a fixed stimulus, a recorded response, a scoring procedure, which is to say a test. And the format has now been running long enough to be judged the way tests are judged, on evidence.

The format's appeal is a standardization appeal, and it is genuine. An asynchronous video interview puts identical questions to every candidate, at any hour, in any time zone, and returns scores without a scheduling bottleneck or a panel's competing moods. That is the discipline interview research has urged on hiring for decades, delivered by default, at a volume no panel could sit through. The trouble is that the category arrived as software before it arrived as measurement, and the space between those two identities is where the risk sits.

The stakes compound with volume. In campus and high-volume hiring, where AI video interviews are now routine, a single configuration decision repeats across every applicant it screens, a property tests have always had and conversations never did. When an interviewer misjudges, the error is local and diluted by the next interviewer. When a scoring model is misconfigured, the error is systemic, identical, and silent, reproduced in every recording it reads.

The credible evidence draws three lines. The usable signal lives mostly in what candidates say, the same place structured human interviews find theirs. The medium itself moves the numbers, so mixing formats across candidates breaks comparability before an algorithm ever runs. And the public evidence behind vendor claims has lagged those claims badly enough that regulators, and at least one prominent vendor, have already acted. Each line ends in an operating rule, and none of the rules is hostile to the technology.

The signal lives in what candidates say

The academic record starts with Hickman, Bosch, Ng, Saef, Tay, and Woo (2022), published in the Journal of Applied Psychology: three samples of mock video interviews totaling 1,073 people, plus a separate 99-person sample retested to check score stability. The researchers trained the scoring models themselves instead of auditing a vendor's sealed product, which is what makes the study legible: the design choices a vendor keeps private are, here, in print.

Two findings order everything else. First, the training target mattered. Models trained to reproduce interviewer judgments of candidates showed stronger validity evidence than models trained to reproduce candidates' descriptions of themselves. A scoring model has to be taught what a good answer looks like from someone's ratings, and the choice of whose turns out to be a validity decision in its own right, made before a single candidate presses record.

A model of this kind is easier to evaluate once it is named for what it is. Trained on interviewer judgments, an AVI scorer is an automated panel member: it learns to anticipate what a structured panel would have said about a recording, and its ceiling is the quality of the judgments it studied, rater quality frozen into software. Trained on self-reports, it becomes a different instrument, a predictor of how candidates describe themselves, which is a different construct with a different relationship to the job. Neither target is dishonest, but only one of them is the judgment hiring already trusts, and the study's validity evidence followed that trust.

Second, among the predictors that carried useful signal, verbal behavior dominated. What candidates said (word choice, the content of the answer, its length) outweighed the paraverbal channel of tone and delivery and the nonverbal channel of face and gesture. Figure 1 sketches that hierarchy. The camera, the component that gives the format its name and most of its unease, was never where the dominant signal lived.

Verbal dominance has a practical edge for anyone configuring the format. If the signal is in the words, the transcript is the instrument's true input, and the questions that elicit codable words, behavioral prompts with a describable structure, matter more than production values on either side of the lens. It also means the format's most contested components, gaze, expression, background, carry a weight the evidence does not endorse.

The study tempers the enthusiasm it invites. Test–retest reliability, the stability of a score when the same person is measured again, was mixed across traits, so an automated score has to be defended trait by trait; "the model is validated" is not a sentence that carries meaning at the grain a decision needs. And the interviewer-trained scores predicted academic outcomes, real signal, collected at some distance from the job-performance criterion an employer buys the format to reach.

Read psychometrically, the pattern is almost reassuring. The machine succeeds where it reads the interview's content, and content is what carries a human interview's validity too: free-form conversation predicts weakly, while structure, meaning fixed questions scored against anchored rubrics, raises the format's record substantially, an evidence trail the series walks in its briefing on unstructured interview validity. On this evidence, algorithmic interview scoring works when it reads speech the way a disciplined interviewer would, and the interview's oldest lesson survives automation intact: the value sits in the structure and the content, whoever or whatever does the scoring.

Mediated interviews rate lower before any model runs

Set the algorithm aside for a moment, because the medium carries its own effect. Blacksmith, Willford, and Behrend (2016) meta-analyzed the comparison in Personnel Assessment and Decisions, pooling 12 studies and 13 samples covering 1,557 people. Interviews conducted through technology drew lower interviewer ratings than face-to-face interviews, a standardized difference of d = −.41, roughly four-tenths of a standard deviation; applicant reactions to the process were also less favorable, at d = −.36. Figure 2 plots both against the face-to-face baseline.

Those comparisons largely predate the current generation of model-scored interviews, and the caution transfers rather than expires: the medium is part of the instrument. Camera framing, lighting, a prompt on a screen, silence where a nod would be: each is a testing condition, and unstandardized testing conditions contaminate whatever is scored downstream. A candidate who interviews poorly by video has not necessarily interviewed poorly; the condition is part of the score.

Standardizing those conditions is cheap by the measure of most measurement problems: identical recording guidance for every candidate, device expectations, framing, retake policy, preparation minutes stated in advance. These are instrument settings, and stating them up front is standardization and candidate courtesy at once.

The authors draw the operational consequences themselves. Never mix media inside one requisition's pool: if some finalists meet a panel while others answer a camera, the two groups' ratings differ for reasons that have nothing to do with the candidates, a standardization failure, not an algorithmic one, and entirely within the hiring organization's control. And findings or norms built on one medium do not transfer to another unexamined. A rating distribution calibrated on face-to-face conversation describes face-to-face conversation; the compared-with-whom logic that governs every norm, treated in full in the series' briefing on what test scores mean, applies to the medium as it applies to the norm group.

The reactions estimate is easy to skip past and expensive to ignore. Candidates rated mediated processes less favorably, and reactions are part of what a hiring process produces: they shape how an organization is perceived by the many people it evaluates and does not hire. The series treats candidate reactions as a measurable outcome in its briefing on candidate experience; the narrow point here is comparative. An organization adopting the format starts from behind on goodwill and should design the rest of the experience, instructions, preparation time, human contact around the recording, to pay that debt down.

Public claims ran ahead of public evidence

Against that academic record stands a market that grew faster than its documentation, and the field eventually checked. Raghavan, Barocas, Kleinberg, and Levy (2020), writing in the ACM's conference on fairness, accountability, and transparency, surveyed the public claims and practices of vendors of algorithmic pre-employment assessment. They found little public documentation of validation practices: predictive claims whose supporting evidence was mostly nowhere a buyer could read it. Where fairness claims appeared, they were frequently anchored to compliance with the four-fifths rule, the adverse-impact screening threshold the series unpacks in its briefing on the four-fifths rule, a litigation floor standing where measurement evidence should stand.

The gap has a mundane mechanism. A predictive claim can be printed the day a product ships, while the evidence for it accrues at the pace of hiring cycles and criterion data; add trade-secret caution about disclosing features and training targets, and the equilibrium is the one the survey documented, strong claims over thin public records. Bad faith is unnecessary for that outcome. It arrives whenever no one is obligated to publish. Figure 3 traces the chain a vendor claim rests on, with the links marked where the survey found the record thin.

Then a correction arrived from inside the market. HireVue, a prominent vendor of automated video interviews, commissioned an independent algorithmic audit from O'Neil Risk Consulting & Algorithmic Auditing and announced in January 2021 that it would discontinue facial-analysis inputs in its assessments (O'Neil Risk Consulting & Algorithmic Auditing, 2021). That record is a practitioner document rather than peer-reviewed research, and it deserves a neutral reading: a vendor examined its own instrument under outside scrutiny and retired an input channel. But the channel it retired is the camera's own, the one the academic evidence never found dominant, and the convergence between the market's retreat and Figure 1's hierarchy is hard to miss.

The shape of the episode matters as much as its outcome: outside examination, a public description, a change to the instrument. That sequence, examination to disclosure to revision, is what the rest of the market has yet to routinize.

Regulation has read the record the same way. Illinois requires notice and consent before AI analysis of a video interview under its Artificial Intelligence Video Interview Act, and New York City's Local Law 144 conditions the use of automated employment decision tools on independent bias audits and candidate notice; the series maps the statute book in its briefing on AI hiring regulations. The premise the laws share is the one this article has been assembling: a scored interview is a measurement procedure, and measurement procedures owe evidence to the people they score. The lesson of the vendor-audit era is narrower than fraud and more useful than outrage: the burden of proof sits with the scorer, and a market left to self-describe will not volunteer it.

Building defensible AI video interviews

A defensible deployment starts from the recognition that the asynchronous video interview is a design space, not a single instrument. Lukacik, Bourdage, and Roulin (2022), in a conceptual review in Human Resource Management Review, map the choices that make one AVI a different measure from another: question format and wording, how much preparation time candidates receive, whether responses can be re-recorded and how many times, how long the response window runs. Each choice plausibly shifts what is being measured. Unlimited retakes sample a candidate's best rehearsed performance; a single take with a short preparation window samples something closer to spontaneous speech. Neither is wrong, but an organization has to know which instrument it configured and defend the scores accordingly.

From there the format inherits the structured interview's checklist, unchanged by the automation: identical questions in a fixed order, uniform time limits, scoring against behaviorally anchored rubrics, mechanics the series details in its briefing on running structured interviews. Whether the scorer is a trained panel or a trained model, the interview's validity is built at design time, in the stimulus and the rubric.

Scoring targets should be constructs the organization can name: communication under pressure, conscientiousness, job knowledge, whatever the role analysis says a recorded answer can carry. The training-target finding is the technical version of this demand, since a model reproduces whatever it was pointed at, and pointing it at a defensible judgment is the decision that carries the rest. The evidence obligations, meanwhile, do not soften because the scorer is software: the field's joint testing standards hold any scoring procedure, whatever its technology, to documented reliability, validity evidence, and interpretable reporting (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014). What that looks like when a platform publishes it is laid out in the science behind 5Profiler.

There is a plain benchmark for that obligation, because older instruments already meet it. A traditional published test arrives with a technical manual: how the instrument was built, what its scores mean, the reliability and validity evidence behind them, the populations it was normed on. The substance is as writable for a neural scoring pipeline as for a paper test; what changes is the chapter on how scores are produced. An automated video interview sold on predictive claims deserves the manual those claims imply.

The remaining requirements belong to candidates. They are owed transparency about what is measured and how, the duty the series examines in its briefing on algorithmic transparency and explainability, and they are owed a human lane: a person who can inspect an anomalous score, hear an appeal, and overrule the model. The lane has to be real, staffed, and reachable from the rejection email, or it is a sentence in a privacy policy; a recording scored by software and reviewable by no one fails candidates at measurement and again at recourse. One further design fact belongs on the list: the candidate's side of the camera is changing too, as applicants bring generative tools into interview preparation and delivery, a shift whose prevalence and handling the series covers in its briefing on candidates using AI.

And one deployment sidesteps most of the evidence burden while keeping the standardization gains: use the measurement to prepare the interviewer rather than to replace them. Structured assessment before the conversation; human judgment inside it. A model that never scores the interview, and instead equips the person conducting it, carries none of the scoring model's evidence debt and all of the structure's advantages.

The score belongs inside a battery

For organizations running or evaluating AI video interviews, three questions do most of the diligence, because each attaches to a finding above.

  1. What was the model trained to reproduce? Interviewer judgments and candidate self-reports are different targets with different evidential fates, and the published evidence favors the interviewer-trained kind (Hickman et al., 2022). A vendor unable to answer has answered.
  2. Which channels drive the scores? If the weight sits on face and gesture, the product leans on the channel the academic evidence never found dominant and the market's own audit history has already retreated from.
  3. What published validity and fairness evidence exists? A compliance line drawn at the legal minimum is a floor; the standards ask for documented reliability, validity, and interpretation (American Educational Research Association et al., 2014), and public documentation is exactly what the vendor survey found missing (Raghavan et al., 2020).

Then standardize the medium. One requisition, one format, one set of testing conditions: if the process is asynchronous video, it is asynchronous video for everyone, and ratings gathered across different media never meet in a single ranking (Blacksmith, Willford, & Behrend, 2016).

There is also a defensible order of operations for adoption. Run the format first as structure without automation, human raters on recorded answers, which captures most of the standardization gain immediately. Introduce model scores as advisory, shadowing human judgments where the two can be compared. Let the model gate an outcome only when the organization has seen evidence it would be willing to show a candidate, and keep the human lane open after that. Every step is reversible, and every step generates precisely the record the diligence questions above ask a vendor to produce.

The score's position comes last and decides the most. An automated video interview score, produced by structured prompts and a validated target, is one structured signal about a candidate. It belongs inside a battery, alongside instruments with longer evidence trails, where the composite logic of the series' flagship briefing on what predicts job performance does its work; how a battery is assembled around a role's profile is the substance of the platform overview. The same demand travels to every AI-era format, including the game-based assessments whose evidence base the series weighs in its briefing on gamified assessments. What the score must never be is the verdict.

The asynchronous video interview will keep spreading; its logistics are too useful and its standardization gains are real. What the evidence adds is a job description for whoever operates it. Score the construct, not the camera. Demand interview-grade evidence for interview-grade claims. And treat every recorded answer as what it has become: an item on a test that someone, named and accountable, is scoring.

Where 5Profiler stands

5Profiler runs the measurement before the conversation and keeps the conversation human. Each candidate's assessment profile generates per-candidate interview probes, the questions worth pressing because the evidence left them open, so the interview starts where the assessment stopped and spends its minutes where only a conversation can go. Interviewers arrive with structure in hand; candidates face a person asking questions their own results made relevant. Where distance puts identity and integrity in question, four proctoring tiers span the range from focus checks to full lockdown. And the room this article opened in gets its population back: a camera on verification duty, a candidate, and an interviewer back in the room.

Read the science behind the platform · See it on your roles

References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
  2. Blacksmith, N., Willford, J. C., & Behrend, T. S. (2016). Technology in the employment interview: A meta-analysis and future research agenda. Personnel Assessment and Decisions, 2(1), 12–20.
  3. Hickman, L., Bosch, N., Ng, V., Saef, R., Tay, L., & Woo, S. E. (2022). Automated video interview personality assessments: Reliability, validity, and generalizability investigations. Journal of Applied Psychology, 107(8), 1323–1351.
  4. Lukacik, E.-R., Bourdage, J. S., & Roulin, N. (2022). Into the void: A conceptual model and research agenda for the design and use of asynchronous video interviews. Human Resource Management Review, 32(1), 100789.
  5. O'Neil Risk Consulting & Algorithmic Auditing. (2021). Description of algorithmic audit: HireVue pre-built assessments. orcaa.com.
  6. Raghavan, M., Barocas, S., Kleinberg, J., & Levy, K. (2020). Mitigating bias in algorithmic hiring: Evaluating claims and practices. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAT* '20), 469–481.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.