Research Integrity & proctoring

Assessment in the age of ChatGPT.

A frontier model scores near the top of a bar exam its predecessor failed. The formats that collapse under that capability, the designs that survive it, and why betting on detection is the wrong response.

On a simulated uniform bar exam, GPT-3.5 scored around the bottom 10% of human test takers; GPT-4, its immediate successor, scored around the top 10% (OpenAI, 2023a). Hold onto that distance between two versions of one product. The jump is usually narrated as a story about AI cheating on assessments: a powerful new tool in dishonest hands. It reads more accurately as a stress test of assessment design, because the exam had not changed and neither had the people sitting it. A format built on transmissible questions had met a machine built to answer them.

Large language models did not invent cheating. What they changed is narrower and more consequential: they collapsed the marginal cost of producing a correct answer to any question that can be carried out of the room and typed back in. Every assessment format now sits somewhere on a single axis, from fully transmissible (the question can leave, the answer can return) to fully embodied (the evidence must be produced by a particular person, observed, in real time). The panic treats that axis as a cliff every format fell off together. The inventory below is calmer: some formats are dead or dying, others survive for structural reasons rather than lucky ones, and the difference is visible in advance.

Getting the inventory right has become urgent because many organizations are converging on the worst available combination: keeping static question banks, the format most exposed, while buying AI-detection tools, a countermeasure whose troubled record is examined below. That pairing preserves the broken design and funds the doubtful remedy. This article works through what actually changed, which formats fail and why, which design properties survive, and why detection is the wrong place to spend an integrity budget.

The price of a correct answer collapsed

Start with the evidence on record, which predates the current model cycle and is sufficient for the argument. OpenAI's technical report on GPT-4 (2023a), the primary public document of the capability jump, reported that on a simulated uniform bar exam the model scored around the top 10% of test takers, where its predecessor GPT-3.5 had scored around the bottom 10%. The finding did not remain a vendor claim: Katz and colleagues (2024) documented the bar-exam result independently, in a peer-reviewed journal. Figure 1 draws the two positions, and the geometry is the argument. An instrument refined over decades to separate prepared candidates from unprepared ones placed one release of a product among the weakest test takers and the next release among the strongest.

Medicine supplies the corroborating anchor. Kung and colleagues (2023), in PLOS Digital Health, found that ChatGPT performed at or near the passing threshold across all three steps of the United States Medical Licensing Examination, without specialized training on the exam. For anyone who runs hiring tests, the detail that matters is the final clause. Nothing was tuned and nothing was studied for; a general-purpose text engine reached the boundary of medical licensure as a side effect of being good at answering transmitted questions.

Two disciplines about this evidence keep the rest of the article defensible. First, these are 2023 results, and they are cited here as a floor: no claim in this article depends on what any current model can or cannot do, and none is made. If newer systems perform better on such exams, the design argument below only sharpens; if they perform worse, nothing in it relies on them. Second, the results say nothing about candidates' character. The bar exam was defeated not by dishonesty but by a property of its own format, and that property, transmissibility, is the one variable an assessment owner fully controls.

Consider the economics before and after. Defeating a serious assessment once required a confederate with real expertise, a purchased proxy, or a leaked answer key, each of which cost money, carried risk, and took time to arrange. The pre-LLM literature had already named the underlying vulnerability: Cizek (1999) documented that an exposed item compromises a form for every future test taker, a finding from the paper era that circulation networks merely accelerated. What the model removes is the need for prior exposure at all. A static question crosses the room's boundary at the moment it is asked, and a fluent answer crosses back seconds later. That is the industrialization: not a new kind of dishonesty, but the collapse of the price of the oldest kind.

AI cheating on assessments tracks the transmissibility of the format

Rank formats by two questions and the pattern organizes itself: can the question leave the room, and can the answer come back in. Both are properties of delivery, not of content quality; a brilliant item in a transmissible format is a brilliant item on a countdown. Figure 2 runs the two questions across seven common formats, and hiring tests appear on every row of the table. The verdicts follow mechanically.

The top of the table is the graveyard. A static multiple-choice bank fails both tests at once: its items circulate to future candidates, the exposure economics examined in the companion article on adaptive testing and item security, and each item can now also be answered live by anyone with a second device. The same verdict covers any static knowledge test, whatever its response format, because recall of transmissible facts is precisely the work a text engine automates first. Unsupervised essays and take-home tasks fail differently. The question hardly matters, because the model performs the work itself, and the submitted artifact certifies possession of an answer rather than production of one. Asynchronous video interviews built on known or predictable prompts sit one rung up, degraded rather than dead: the delivery is human, but the content can be generated in advance and rehearsed, so the format increasingly measures presentation of an answer rather than the thinking behind it.

Each of these failure modes predates the technology that industrialized it. Tippins and colleagues (2006), writing about unproctored internet testing long before language models existed, identified precisely this pair of risks, exposure of content and assistance during the sitting, and pointed to supervision and verification as the answers; the fuller account of that debate and of the two-stage verification design it produced is in the companion article on unproctored testing and verification. What the field lacked in 2006 was urgency, because assistance was bounded by whoever a candidate could recruit and pay. It is now bounded by nothing a candidate cannot afford, which is why during-sitting AI assistance has taken its place in the standing fraud-vector taxonomy mapped in the companion article on remote assessment fraud. The vectors are that article's territory; the design response is this one's.

What survives is designed around the person, not the artifact

The surviving rows of Figure 2 share no subject matter and no technology. What they share is a set of design properties, each of which removes one link from the transmission chain. Stated as properties rather than products, they can be checked against any format, including ones this article does not list.

The first is supervision of the sitting. When delivery is supervised, in a room or by live remote observation, the answer must be produced by the enrolled person, at the appointed time, under watch. This is a structural property, not a psychological one: whatever a candidate's intentions, an outside answer has no unobserved route back to the response. The behavioral science of monitoring, what visible observation does to the decision to attempt anything at all, belongs to the companion article on proctoring as deterrence rather than detection; supervision matters here simply as the condition under which the transmission channel closes.

The second is per-candidate item sequences. A question that no other candidate will ever be asked is a question not worth harvesting: there is no stable form to photograph and no answer key with resale value, so the incentive to exfiltrate content collapses along with its payoff. The machinery that makes this practical at scale, calibrated banks, exposure control, adaptive assembly, is the subject of the adaptive testing article and is not re-argued here. The property this argument needs is only that each sitting is unique to the person in it, which converts content theft from an investment into a waste of effort.

The third is observed process rather than submitted artifact. A take-home certifies that work arrived; an observed work sample certifies that work happened. When the assessor can see the making, the sequence of attempts, corrections, dead ends, and recoveries, the evidence is the process itself, and a process is not transmissible the way a deliverable is. This is why organizations moving stakes off take-homes tend to land on supervised work samples rather than on longer or cleverer take-homes: the improvement is in what gets witnessed, not in what gets asked.

The fourth is live interaction with follow-ups. A structured interview, the same job-related questions in the same order with anchored scoring, adds a property no static format has: the next probe depends on the last answer, so the exchange cannot be fully rehearsed in advance. Under supervision, the person in the seat is the person answering, and each follow-up is generated in the moment rather than drawn from anything that could have circulated. The claim being made is deliberately narrow. It is not a claim that models are incapable of conversational performance; the 2023 record cited here says nothing either way about that frontier. The claim is that a supervised live exchange is an arrangement in which assistance has no unobserved channel, and arrangements, unlike capability frontiers, do not move on a release schedule.

Survival, throughout, is conditional. Remove the supervision from the sequenced test and it rejoins the graveyard; publish the interview's probes and liveness degrades toward the scripted-video row. There is no LLM-proof assessment, only formats that leave less in transit and the discipline of keeping them that way. That humility is itself the design argument: a property you must maintain is at least a property you can inspect, while a detector's error rate is somebody else's secret.

Detection is a bet the evidence has already graded

The market's favored alternative asks for no redesign at all. Keep the essay, keep the take-home, keep the bank, and buy a classifier that inspects submitted text and estimates whether a machine wrote it. The appeal is plain: detection promises integrity as a procurement line item rather than a design change. The reliability record is where the promise fails.

The record here is short and unambiguous. OpenAI released a classifier intended to detect AI-written text and withdrew it within the same year, citing its low rate of accuracy (OpenAI, 2023b). The maker of the generator itself, with full visibility into how the text is produced, judged its own detector unfit to keep in service. That is one data point, and it should not be inflated into a verdict on every tool. But a buyer weighing AI detection reliability should notice what kind of data point it is — the generator's own maker's attempt, retired for the precise deficiency the product existed to eliminate.

The structural argument explains why the burden of proof belongs on the detector. Generator and detector are trained on the same substrate, and the generator's entire objective is to produce text with the statistical signature of human writing; every improvement on one side is, by construction, a harder problem for the other. Output detection is a contest against an opponent whose training goal is your false-negative rate. Contests of that shape can swing either way at any moment, which is exactly what makes them unfit foundations for decisions that must still be defensible months later, in front of a rejected candidate or a court.

There is also a cost asymmetry on the far side of the error. A detector that is wrong about a candidate has not merely missed a cheat; it has manufactured an accusation against an innocent person, and what organizations owe the people their flags touch, in evidence, review, and a route to reply, is the subject of the companion article on integrity flags and false positives. Figure 3 sets the two spending paths side by side. Lane A preserves the exposed format and enters an open-ended contest that began from a documented failure. Lane B changes the design so that the question detection was hired to answer, whether a machine produced this artifact, stops arising, because the sitting produces no unattended artifact in the first place.

Both sides of the hiring desk now hold the same tools

Candidates reading integrity policies may fairly observe that the tools run in both directions: employers screen applications with algorithms and generate outreach at scale, then restrict candidates' use of the same technology. The symmetry deserves resolution rather than resentment, and the resolution is one design principle applied to both sides of the desk.

Assistance is legitimate where it is disclosed and where the assessment measures output a person would produce with tools on the job; it is a threat to measurement where the assessment targets personal capability and the assistance substitutes for it. The obligation that follows is likewise symmetrical. State in the invitation which sittings permit which tools and how the organization itself uses automation in screening, then build the restricted sittings so the restriction is structural rather than honor-based, because a rule the design does not enforce is a request.

An organization that automates its own side of the desk while leaving its rules implicit has not been cheated when candidates improvise; it has declined to specify the game.

Audit for transmissibility, then move the stakes

The practical response to AI cheating on assessments compresses into an audit any assessment owner can run this quarter, and Figure 4 lays it out as a flow. For every assessment currently carrying stakes, ask three gate questions in order: whether the item pool is static and reused across candidates, whether the questions could be known before the sitting through circulation, prediction, or simple reuse, and whether the answer could be produced outside the sitting by anyone or anything other than the candidate. A yes at any gate is a finding, and each finding has a specific remedy rather than a clever one.

Retire or rotate what fails the first two gates. A static pool's value decays with every sitting, so content review on a calendar, retirement of exposed items, and, where volume justifies it, migration to per-candidate assembly are maintenance, not luxury. Formats that fail the third gate need their stakes moved rather than their wording improved: an unsupervised essay can continue as a low-stakes exercise or as a writing sample discussed live, but decision weight belongs on supervised, sequenced, and observed formats, and the delivery platform is where those properties are either engineered or absent. Essays and take-homes keep a place as conversation starters and practice material. The offer decision is what moves.

Publish the AI policy in the candidate invitation. Ambiguity manufactures accidental cheats: a candidate who polishes a take-home with a model in the absence of a stated rule has not clearly broken one, and an integrity case built on an unstated rule collapses on first contact with an appeal. The invitation should say which tools are permitted at which stage, which stages are supervised and why, and how the organization itself uses automation in screening. Candidates read design. A stated policy attached to a supervised, per-person format signals seriousness more credibly than any warning banner stapled to a static form.

Spend nothing on output detectors where the call is high-stakes. Given the detector's own history, the narrower point is decisive on its own: a control of unestablished reliability cannot carry a decision that must survive challenge. If an organization runs a detector at all, its output belongs at the start of a human review with evidence attached, never at the end of an automated rejection, the discipline the integrity-flags article develops in full. Delivery integrity and measurement quality remain separate purchases even then: a perfectly defended sitting still needs instruments worth defending, which is the science behind the platform's own burden to meet.

The reframe this article argues for ends as an inversion of blame. When a general-purpose text engine scores among the strongest test takers on a licensure exam, as the 2023 results on record show one did, the exam has learned something about itself: its questions travel, and its answers travel back. Formats with that property were failing quietly long before the technology arrived to fail them loudly, because transmissibility was always the same door through which leaked forms and coached answers came in. AI cheating on assessments will keep the panic headlines; assessment owners hold the calmer lever. Close the room by construction — supervise the sitting, sequence per candidate, observe the work, converse live — and what remains for a model to do is what remains for any outsider to do, which is nothing with an unobserved route to the score.

Where 5Profiler stands

5Profiler puts its weight where this evidence points: on delivery design rather than output detection. Sittings are supervised live, and every candidate receives a different, adaptively assembled sequence, so there is no stable form to harvest and no transmissible answer for a tool to supply in advance. A leak from any single sitting captures only that candidate's own sequence, content no future candidate will ever see, so the harvest has no buyer and the effort no payoff. That is the platform's standing answer to the two questions that organize this article: the question cannot usefully leave the room, and no outside answer has an unobserved route back in.

Read the science behind the platform · See it on your roles

References

  1. Cizek, G. J. (1999). Cheating on tests: How to do it, detect it, and prevent it. Mahwah, NJ: Erlbaum.
  2. Katz, D. M., Bommarito, M. J., Gao, S., & Arredondo, P. (2024). GPT-4 passes the bar exam. Philosophical Transactions of the Royal Society A, 382(2270), 20230254.
  3. Kung, T. H., Cheatham, M., Medenilla, A., Sillos, C., De Leon, L., Elepaño, C., … Tseng, V. (2023). Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digital Health, 2(2), e0000198.
  4. OpenAI. (2023a). GPT-4 technical report. arXiv:2303.08774.
  5. OpenAI. (2023b). New AI classifier for indicating AI-written text [Product announcement; updated July 2023 to note discontinuation].
  6. Tippins, N. T., Beaty, J., Drasgow, F., Gibson, W. M., Pearlman, K., Segall, D. O., & Shepherd, W. (2006). Unproctored internet testing in employment settings. Personnel Psychology, 59(1), 189–225.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.