Research The AI era

The candidate’s other window.

By 2023, 17% of surveyed early-career candidates had already used generative AI in applications and 70% expected to — while GPT-4 outscored 98.8% of humans on verbal reasoning. The policy spectrum, priced honestly.

The most consequential new participant in the hiring funnel is not a vendor, a regulator, or a rival employer. It is the second window open on the candidate's screen. Candidates using AI now draft the cover letter, rehearse the interview, and sit beside the assessment, much as an earlier generation used spell-check and a search engine: as ordinary preparation, mostly at stages where nothing told them otherwise. This is a change in normal candidate behavior, and reading it as a scandal is the fastest way to get the response wrong.

Draw the boundary first. This is not the impersonation problem: proxy test takers and contract cheating services, the industrialized fraud catalogued in the companion briefing on remote assessment fraud, put a different person behind the screen. The candidate in the second-window case is exactly who they claim to be and is doing the work themselves, with assistance capable enough to change what the work shows. Impersonation attacks identity; assistance changes conditions; the two problems need different answers.

The argument of this briefing is that the answer is a published rule, stage by stage, and that the employer owes it. A funnel that has published no candidate AI policy has not actually banned anything. It has left the call to each candidate's private judgment, with no visibility into how the call lands. The design work runs through three moves: decide, for each stage, what construct the stage measures and whether assistance changes it; tell candidates which uses are allowed and disallowed there; and anchor the verdicts that matter in conditions where the measurement survives. The rest of the article is the evidence that funnels need this now, and the detail of how to build it.

Candidates using AI stopped being the exception in 2023

The core prevalence evidence dates from 2023, early in the tool cycle, and its provenance matters as much as its values. That year the assessment firm Arctic Shores commissioned survey research with the polling company Opinium, asking 2,000 students and early-career candidates about generative AI in their job applications and assessments (Arctic Shores, 2023). In that sample, 17% had already used generative AI in an application or an assessment, and 70% expected to use it within the next 12 months. This is vendor-commissioned research, from a firm with a commercial stake in the answer, and it should be read with the caution that provenance calls for. It cannot be read as marginal.

Two later practitioner readings corroborate the picture. A 2024 practitioner survey by ResumeTemplates.com found that 22% of Gen Z applicants had used ChatGPT to create a résumé or cover letter as of May 2024 (ResumeTemplates.com, 2024). Kickresume's platform telemetry counted more than a million job seekers using AI inside its résumé tooling across 2025 (Kickresume, 2025). These are practitioner readings, and this article treats them that way: firms positioned to observe the behavior directly, reporting what they saw, without peer review. What the readings establish is the direction and the floor. Figure 1 charts the licensed shares with their sources attached; the bars share one proportional scale, and each survey keeps its own panel.

This series usually stands on peer-reviewed ground, and this section deliberately does not, because the prevalence of a behavior this new is not yet in the journals; the parties positioned to measure it early are the vendors and platforms that sit inside the funnel. Their numbers are used here with the labels on: commissioned by interested parties, unaudited, and directionally consistent across sources with different interests. The argument that follows leans only on the property the readings share, which is their direction, and on a floor that no serious observer of the tool cycle disputes.

The two Arctic Shores bars read as a pair with a time axis inside them. The 17% is behavior already performed by 2023; the 70% is intent declared alongside it, the direction of travel measured before most of the travel had happened. The survey's age is exactly its use: it caught the norm while the norm was forming. Nothing in the tool cycle since, in capability, distribution, or workplace adoption, has pushed toward less use, and there is no serious reading on which candidate use today sits below the intent measured then. The floor is established. Only the ceiling is open.

One more feature of the prevalence picture changes how it should be read: very little of this use broke any stated rule. Most application processes said nothing about generative AI at any stage in 2023, and most still say nothing at most stages. A candidate who drafts a cover letter with a model in a funnel that stated no rule has broken no rule. The assumption that application work would be done alone was real, but it was never written anywhere a candidate could read it. The integrity framing therefore starts in the wrong place: before any question about candidate conduct comes a question about employer specification, of what the process told candidates, stage by stage, about the conditions of the work.

Assistance strong enough to change what a score means

Prevalence justifies writing a policy; capability is what makes the policy urgent. The Arctic Shores program included the firm's own testing of what the assistance can do: run against verbal reasoning items, GPT-4 scored above 98.8% of human candidates (Arctic Shores, 2023). That is the vendor's own figure, produced on its own items, and the argument needs no more precision than that, because it does not turn on the decimal. Assistance at that level does not nudge an unproctored score; it replaces the thing the score was supposed to be about.

Measurement theory has exact language for what follows. A test score carries meaning under stated administration conditions. The Standards for educational and psychological testing (American Educational Research Association et al., 2014) treat those conditions as part of what a score means, and treat changes in them as changes in what may be inferred and in whether scores remain comparable across candidates; the SIOP Principles carry the point into hiring, where standardized administration is a condition of validity (Society for Industrial and Organizational Psychology, 2018). When a second window may or may not have been open, an unproctored verbal score stops meaning "this candidate's reasoning" and starts meaning "the output of this candidate plus whatever the other window held." The construct shifts silently, candidate by candidate, in proportion to who used what.

Comparability is the sharper casualty. Two candidates with identical unproctored scores may have produced them under different effective conditions, one alone, one assisted, and the score report shows no trace of the difference. A ranking built on those scores treats numbers gathered under different conditions as one measurement, which is exactly the practice the comparability standards exist to prevent. AI in the hiring funnel is best understood at this altitude: as an uncontrolled administration condition, the kind of problem measurement people have known how to think about for decades.

The operational version of the loss is easy to picture. A talent team benchmarks this year's graduate intake against last year's on an unproctored reasoning screen and reads the rise in mean scores as a stronger cohort. Under the prevalence evidence, part of that rise is a change in administration conditions misread as a change in talent. Norms drift, cut scores calibrated on unassisted cohorts begin to misclassify, and every downstream decision inherits the confusion. No one made an error at any step; the conditions changed underneath the instrument.

Exposure is not uniform across the funnel, and the non-uniformity is what design can use. At one pole sit written application materials: producing fluent prose from a prompt is precisely what the tools do, which made ChatGPT-drafted job applications the first visible form of the behavior; unproctored online tests share that pole, and the two-stage designs that re-verify an unproctored screen under supervision, mapped in the companion briefing on unproctored testing and verification, exist for exactly this exposure. At the other pole sit live, identity-verified stages and interactive work samples, where assistance would have to operate in real time, under observation, inside a task that responds to the candidate. Figure 2 arranges the funnel's stages between those poles; which item formats collapse under a capable model and which survive is mapped in the briefing on AI-resistant assessment design, and the short version is that exposure tracks how transmissible the task is.

The traffic also runs the other way: while candidates draft with models, employers increasingly hand the resulting text to models of their own to score, and the companion briefing on LLM scoring of open responses examines that side of the exchange. The compounding case is the one to hold on to. A funnel in which model-written answers are model-graded, with neither side having said so, is a funnel in which nobody can state what was measured, for any candidate, at any stage. That prospect, more than any single inflated score, is the case for making conditions explicit everywhere.

A ban selects for non-disclosure

The reflexive policy response is prohibition: a line in the application portal declaring that AI use is forbidden. It reads as decisive, costs nothing to write, and is spreading accordingly. It fails on three fronts.

The first front is enforcement. A ban is only as strong as the means of detecting violations, and the detection evidence, reviewed in the AI-resistant design briefing and not re-argued here, does not support building policy on it. An unenforceable rule still has effects. They are just not the intended ones.

The second front is who the unenforceable rule filters. Candidates who read the ban and take it seriously comply, disclose, or withdraw. Candidates who do not take it seriously proceed and say nothing. The ban does not remove AI from the funnel; it removes disclosure from it, and the candidates it screens out are disproportionately the ones who told the truth. The fault is structural, and worth being exact about: the selection pressure comes from the gap between what the rule forbids and what the process can see, and it operates on honest candidates regardless of anyone's intent. A policy that punishes candor at intake is a strange foundation for an integrity program.

The decision looks different from the candidate's chair. Under a ban, a candidate who has drafted with a model can withdraw, can disclose and accept whatever penalty an unstated process applies, or can proceed and say nothing. The first path removes truthful candidates from the pool. The second is a bet on the goodwill of a process that has just called the behavior cheating. The third is what remains, and it is what the funnel's own design selects for.

The third front is the jobs. Across a widening set of roles, drafting with a model, prompting it well, and editing its output critically is becoming job content. A funnel that prohibits the tool everywhere is assessing candidates under conditions the job will never reproduce, and the mismatch is double: it misses fluency the role rewards, and it rewards unassisted polish the role will not need. For those roles the live design question is the reverse of the ban: whether some stage should require the tool, and grade what the candidate does with it.

The professional literature reached this ground before the current tools did. Tippins, Oswald, and McPhail (2021), in a call-to-action review of AI-based selection published in Personnel Assessment and Decisions, set out the field's scientific, legal, and ethical concerns and the obligations that follow for employers: standardize conditions, document what is measured, and be transparent with the people being measured. Those verbs were written about employers' own algorithms, and they transfer intact to employers' policies on candidates' tools. The stance of this briefing follows from them. Where the rules are silent, candidates using AI are not cheating anyone, and the deficit sits with the party that owns the process. The employer owes the rule; the rest of the argument is about what a good one looks like.

A workable candidate AI policy is set stage by stage

Between prohibition and mandate the policy space is a spectrum, and every position on it trades an enforcement burden against what remains measurable; every position also decides, implicitly, whether candidates using AI are treated as a risk, a resource, or both depending on the stage. Figure 3 sketches the positions and the trade each one makes. The middle of the spectrum repays a closer look; the figure carries the rest.

Disclosure-based positions (candidates report what they used) are the tempting middle and the weakest ground. A disclosure regime produces a paper answer to the construct question while leaving verification exactly where a ban leaves it, on the honor system, and it reproduces the ban's structural unfairness in a politer register: the report is unverifiable, so candidates who report accurately subsidize those who do not.

The defensible ground is restriction by stage, because its logic follows the construct. Stages where the construct is the candidate's own unassisted capability, the reasoning screen, the final verified assessment, are designated AI-free and run under conditions that make the designation true: identity verified, the session observed, the proctoring intensity matched to the stakes. The deterrence evidence, reviewed in the companion briefing on deterrence and detection in proctoring, is encouraging on exactly this point: announced observation prevents most attempts from being made at all, a gentler outcome than catching them afterward. Stages where the construct is the work product, a drafting exercise, an analysis in a role that will use these tools daily, are designated AI-allowed and scored as what they now are: the candidate's output under realistic tool access.

Scoring an AI-allowed stage needs a rubric written for collaboration rather than recall: does the candidate direct the tool or transcribe it, verify its claims or forward its errors, show judgment about what to keep and what to cut? Those behaviors are observable in the artifact and in a short follow-up conversation about it, and they are closer to the day-one reality of tool-heavy roles than any unassisted essay. The rubric also protects the stage from becoming a loophole: allowed does not mean ungraded, and candidates should know that the collaboration itself is being read.

The third component faces the candidate: publish, at every stage, which designation applies and what it means in practice. The applicant-reactions evidence, reviewed in the companion briefing on candidate experience in assessment, is blunt that candidates judge selection processes by whether the rules were knowable and evenly applied, and a published per-stage rule is that respect delivered in procedural form. It converts the integrity question from a trap into an instruction. It also changes what a violation means: once the rule is explicit, use where use is disallowed is a choice against a stated condition, an integrity event the employer can act on cleanly, because no candidate can truthfully say that nobody told them.

The return on the harder policy is measurement that means something. Scores from AI-free stages regain their construct: they describe the person, under stated conditions, comparably across candidates. Scores from AI-allowed stages acquire a different and equally usable meaning: performance under the conditions the job will actually offer. And the assessment file regains its defensibility, because for every number in it the organization can state the conditions under which the number was produced, which is the documentation posture the professional standards have asked of selection programs all along (American Educational Research Association et al., 2014).

Publish the rule before the next cohort applies

The implications are concrete enough to put on a calendar, and the work belongs jointly to talent acquisition, the assessment owner, and counsel, because the rule touches candidate communication, construct decisions, and legal defensibility at once.

  • Write the per-stage policy this quarter, and publish it where candidates will see it. For every stage in the funnel, one line: what the stage measures, what assistance is allowed, what conditions apply. A stage whose line is hard to write is usually a stage whose construct was never decided, and that discovery alone repays the exercise; a policy this concrete is also a configuration an assessment platform can hold natively and enforce per stage.
  • Re-read every unproctored score the funnel currently relies on. Given the prevalence evidence, a standing unproctored reasoning score is a number gathered under unknown conditions. Treat it as a screening signal, and confirm it under observed conditions before it carries verdict weight; the verification designs exist and are documented.
  • Move verdict weight to verified stages. Offers, final-gate rejections, and rankings that decide interviews should rest on scores gathered where identity was checked and conditions were observed. How observation, verification, and lockdown are engineered is a platform security question; where in the funnel they apply is the policy's to say.
  • Measure the policy after it ships. A published rule moves behavior: completion rates, stage conversion, and time-in-stage will shift, and those shifts are data about how candidates read your conditions. An organization that watches them can tune the rule. One that assumes compliance has returned to delegating the decision to the second window.

A published candidate AI policy also works on the funnel from the outside. Candidates read selection processes as previews of the employer, and a process that states its conditions, stage by stage, previews an organization that knows what it is measuring and says so. While stated conditions remain rare, that signal is cheap to send and hard to fake, and it lands on exactly the candidates an assessment-heavy funnel most needs to keep. The second window is not going to close, and the tools behind it will improve on their own schedule, with no reference to anyone's hiring calendar. The organizations that publish their rule get to decide what their scores mean. The ones that publish nothing will keep selecting for non-disclosure, and will keep calling it integrity.

Where 5Profiler stands

Per-stage conditions are native to how 5Profiler runs a funnel. Assessment owners assign one of four proctoring tiers to each stage, from focus-only checks to full lockdown, so an AI-allowed drafting exercise and an AI-free verified sitting can operate in one funnel with conditions matched to what each stage claims to measure. Identity and session integrity are verified on the stages that carry verdicts, and the report shows the conditions each score was produced under, so a reviewer opening it knows which rules applied. Candidates see those rules before the sitting starts, stated in plain language. A funnel configured this way has answered the question this article opened on: at every stage, it says which window may be open.

Read the science behind the platform · See it on your roles

References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
  2. Arctic Shores. (2023). Generative AI and the future of early-careers recruitment (survey research with Opinium, N = 2,000). arcticshores.com.
  3. Kickresume. (2025). AI job search data report. kickresume.com.
  4. ResumeTemplates.com. (2024). Survey: Gen Z job applicants and ChatGPT. resumetemplates.com.
  5. Society for Industrial and Organizational Psychology. (2018). Principles for the validation and use of personnel selection procedures (5th ed.). Industrial and Organizational Psychology, 11(Suppl. 1), 1–97.
  6. Tippins, N. T., Oswald, F. L., & McPhail, S. M. (2021). Scientific, legal, and ethical concerns about AI-based personnel selection tools: A call to action. Personnel Assessment and Decisions, 7(2), 1–22.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.