Research Integrity & proctoring

The flag is not a verdict.

Run a highly accurate detector across thousands of honest candidates and it will still accuse the innocent at scale. The base-rate arithmetic of integrity flags, and the review process that keeps them defensible.

Picture the same integrity flag firing at two employers. At the first, the flagged candidate receives a letter three sentences long: irregularities were detected during your session, your result has been invalidated, this decision is final. At the second, the candidate receives the captured evidence itself, a note that a reviewer will examine the session in context, and an invitation to explain the anomaly before anything is decided. Identical detector, identical threshold, identical anomaly; the only variable is the institution reading the flag.

Proctoring false positives are what separate those letters, and their scale is set by numbers most buyers never multiply. Any detector screening a mostly innocent population will generate accusations of people who deserve none, and depending on prevalence and threshold, the wrongly flagged can rival or outnumber the cheats caught. That is not a defect in one vendor's model. It is a property of detection itself, and it holds no matter how good the software gets.

This article is the piece of the ledger twice deferred by its companion pieces on assessment fraud and proctoring. The companion review of remote assessment fraud built the case for layered detection and noted, near its end, that false positives carry liabilities of their own; the study of proctoring as deterrent finds that monitoring earns most of its value before the attempt, and hands whatever its detectors produce afterward to review. What follows prices that hand-off. The conclusion is not that detection should stop. A flag is the start of a process, never a verdict a dashboard delivers, and organizations that hold that line protect two things at once: candidates from false accusation, and their own decisions from becoming indefensible the first time one is challenged.

The arithmetic nobody runs on their own detector

Three quantities decide what a flag is worth, and only two of them appear in marketing material. Sensitivity, the share of actual cheats a detector flags, is the number that sells. Specificity, the share of innocent candidates the detector correctly leaves alone, is the number that governs how often the alarm cries wolf. The third, positive predictive value, is the probability that any given flag points at real misconduct, and it is the only number a reviewer holding a flagged file actually needs. It cannot be printed on a spec sheet, because it depends on something no vendor controls: prevalence, the share of the screened population that is actually cheating.

The three are locked together by one line of algebra: PPV = p·se / (p·se + (1−p)(1−sp)), where p is prevalence, se sensitivity, and sp specificity. Every value that follows is computed from that formula under stated assumptions; none is a measurement of any product, this platform's included. Hold sensitivity at 80% and suppose 5% of candidates cheat. With specificity at 99%, 81% of flags are true. At 95%, the true share drops to 46%. At 90%, it is 30%. Now suppose the population is cleaner, with 2% cheating, and run the same detector again: 62%, 25%, and 14%.

Figure 1 draws all six results, and the pattern deserves a slow look. A change in specificity that sounds trivial, from 99% to 95%, does more than shave a few points off flag quality: it tips the majority of the queue into error. And prevalence, the variable nobody chooses and nobody measures well, moves the answer as much as the detector does. Under entirely plausible settings, a good detector screening a mostly clean population, most flags are honest candidates.

This is why the number a buyer most needs is the one no proposal contains. Sensitivity and specificity can be estimated in controlled studies, where the ground truth of who cheated is known because the design planted it. Prevalence cannot: it varies by role, stakes, season, and region, and no organization knows its own rate, because measuring it would require the very detection whose quality is in question. Any claimed flag accuracy therefore assumes a prevalence silently, and the assumption matters more than the algorithm. The structure is familiar from medical screening, where a good test for a rare condition still produces mostly false alarms; the difference is that a screening program schedules a confirmatory test by default, and a hiring funnel schedules whatever its workflow says a flag means.

Hold on to the smallest number in the figure. At 90% specificity and 2% prevalence, only 14% of flags mark real misconduct. An organization that treats every flag as a finding is, at those settings, wrong far more often than right, systematically and silently.

Ten thousand candidates, one queue of flags

Percentages hide people, so run a single cohort at full scale, with every value computed from the same stated assumptions: 10,000 candidates, 5% prevalence, 80% sensitivity, 95% specificity. The pool contains 500 cheats, and the detector performs exactly as designed, catching 400 of them. It also misfires across the 9,500 innocent candidates, flagging 475 of them. The queue that reaches a reviewer holds 875 names. Of those, 400, or 46% of the queue, are cheats; the other 475, 54% of it, are candidates whose sittings were clean. Nor is a cohort of 10,000 hypothetical: a single large employer's season of campus hiring runs one, and the queue of 875 names that comes out of it is a workload somebody must own.

Figure 2 shows the composition to scale, and the visual point is the one aggregate reassurances miss: the flag queue does not inherit the population's innocence. The population is 95% clean; the queue is 54% clean. A detector is a machine for concentrating suspicion, and at these settings it concentrates suspicion on more innocent people than guilty ones. Each of those 475 files is a candidacy that ends, an explanation that is never requested, or a review that gets run, depending entirely on what the organization has decided a flag means.

Now give the queue to software instead of a person. An automated integrity score that quietly rejects flagged candidates, or quietly ranks them down, executes the error rate at machine speed with nobody watching it happen. At the settings above, such a system discards more innocent candidates than cheats, and it does so invisibly: the rejected rarely learn why, the employer never learns which, and nothing on a dashboard distinguishes the 400 from the 475. The exposure compounds silently, legal, reputational, and ethical at once, until a single wrongly accused candidate makes it loud. The fraud review called a flag without reviewable evidence a liability pointed in both directions; the cohort above is that liability, counted.

Why proctoring false positives cannot be engineered away

The tempting response is to demand a better detector, and better detectors exist; nothing here argues against building them. But three structural facts guarantee that proctoring false positives survive every upgrade. The first is the threshold trade. Sensitivity and specificity are not independent dials but one dial read from two sides: a detector classifies by comparing evidence against a cutoff, and moving the cutoff to admit fewer false alarms necessarily waves through more real cheats, while moving it to catch more cheats sweeps in more of the innocent. An operator chooses which error to prefer. No operator gets to decline both.

The second is the world the detector watches. Flag models assume an implied normal sitting: one face, steady light, stable bandwidth, a quiet room, eyes on the screen. Real candidates sit assessments in shared apartments where someone crosses the frame, on connections that stutter and drop video mid-section, under backlighting that defeats face tracking, within reach of doorbells and children, and in the lifelong habit of reading questions aloud or staring at the ceiling to think. Each of these is an anomaly to a model and an ordinary evening to a person. These environmental flags also do not land evenly: candidates with disabilities or assistive setups can trip patterns tuned to a narrow default picture of test-taking, an equity dimension this series examines separately and with the care it requires. The benign explanations are not edge cases; they are most of what a detector's errors are made of.

The third is the asymmetry of costs, and it is the fact the whole design should pivot on. A missed cheat corrupts one data point. The damage is bounded, and it is recoverable downstream, because a shortlisted score can be re-verified under supervision before anyone relies on it. A false accusation lands on a person. It can end a candidacy for a reason the candidate never learns, attach a suspicion to a name inside an applicant tracking system, and, wherever it is voiced, follow someone into rooms the employer will never see. One error is a measurement problem; the other is a harm to a specific individual who did what was asked. Figure 3 weighs the two, and the pans do not balance, which is why the response to a flag has to start at nothing and escalate deliberately rather than start at rejection and walk back.

The profession settled this before the software existed

The testing profession has run detection programs since long before webcams, and its literature is unambiguous about what detection output is. Cizek (1999), in a book-length treatment of cheating on tests, states the position plainly: statistical indicators of cheating are grounds for investigation, not conclusions, and corroborating evidence and review are required before action. Wollack and Fremer (2013), whose edited handbook codified test security as a working discipline, are just as direct: security decisions require documented evidence, defined response procedures, and review, because detection alone is not a decision. The International Test Commission's test-security guidelines (International Test Commission, 2014) carry the same requirement to program level: when integrity incidents arise, the response follows documented procedures under defined responsibilities, and detection is one component of a security program rather than a verdict it can issue.

Investigation, in that literature, has a specific shape. A statistical indicator says that a pattern is unusual, and no more than that: it cannot say that a person cheated, because unusual patterns have innocent generators, from lucky guessing to shared preparation to the environmental accidents catalogued above. Corroboration means independent lines of evidence converging on the same account, the captured artifact that shows what actually happened, the opportunity that made misconduct feasible, a result out of keeping with everything else known about the candidate. And when the lines do not converge, the investigation ends and the score stands, an outcome the test-security literature treats as the process working, not failing.

The measurement standards say the same thing from the candidate's side. The Standards for Educational and Psychological Testing, issued jointly by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education (2014), entitle test takers to notice, to procedures for challenging scores, and to decisions based on adequate evidence. Read together, the three sources amount to due process for cheating accusations, written into professional norms decades before proctoring software shipped its first flag. An assessment integrity review that ends candidacies on unexamined output is not a modern efficiency; it is a departure from the standards of the very field whose instruments it uses.

What changed with software is volume. An automated stack can raise more flags in a week than a human proctor raises in a season, which makes the review requirement more load-bearing, not less, because the false share documented in Figure 1 arrives at the same scale as everything else. The professional consensus was built for a world of rare, carefully investigated cases. Automation does not repeal it. It industrializes the input to it.

A process that survives its first challenge

The prescriptions assemble into a specific pipeline, drawn in Figure 4. It begins at capture: a flag that does not carry its evidence, the recording, the event, the timestamps, whatever raised it, cannot be reviewed, only believed. It proceeds through a context check that screens the benign explanations first, connectivity, environment, accessibility, disclosed conditions, before any question of intent is reached. Only then does it arrive in front of a human being with the whole file, and the reviewer's question is not "did the software fire?" but "what does the evidence, in this candidate's context, actually support?" That question has to be answered the same way on a Monday as on a Friday, which is why mature programs define review criteria and train the people who apply them: consistency across cases is itself a fairness requirement, not a bureaucratic taste.

The response that follows is an escalation path, followed only as far as the case demands. Where context explains the signal, the answer is no action and a cleared file. Where the question is whether the result is real, the answer is supervised verification, re-testing the score under observation, the two-stage design detailed in the companion article on unproctored testing; a genuine score survives a retest, and an inflated one does not, which resolves the case on evidence rather than inference. Where uncertainty remains worth recording, the score is annotated rather than the candidate rejected. Only where evidence has hardened does the decision itself come under review. At every step the candidate is notified and heard before anything becomes final, and the whole file, evidence, reasoning, response, and outcome, is kept, because the record is what the process will be judged on later. How capture and retention are engineered is a platform security question; that they exist is a governance one.

This shape is not merely defensive. Gilliland (1993) applied organizational justice theory, the study of what makes procedures feel fair independent of their outcomes, to selection systems, and identified the procedural rules that determine whether candidates accept a process and its results: consistency of administration, accuracy of the information used, and a genuine opportunity to respond. The pipeline in Figure 4 is those rules operationalized. Consistency lives in the defined escalation path, accuracy in the evidence requirement and the context check, voice in notice and reply. A process built this way is harder to challenge and easier to accept at the same time, including for candidates whose flags are ultimately sustained: in Gilliland's framework, people can live with an adverse outcome when the procedure that produced it was visibly sound.

The audience for that acceptance is larger than the flagged. Hausknecht, Day, and Thomas (2004), in a meta-analysis of applicant reactions research, found that how candidates perceive a selection process carries over to how they perceive the organization itself. The treatment of the flagged is visible to the unflagged: cohorts talk, and a process that accuses by template teaches an entire applicant pool what the employer is like when its software is confident and wrong.

Run the flag queue like a decision you will defend

For organizations buying or operating remote assessment, the evidence on proctoring false positives compresses into four operating rules.

  • Never auto-reject on a flag. Wire the constraint into the workflow, not just the policy: no status transition from flagged to rejected without a named reviewer and a recorded reason. If flag volume makes review feel unaffordable, the volume is a message about the threshold, not about the review.
  • Publish the process to candidates. Before the sitting: what is monitored, what happens when something is flagged, how to respond, who decides. Disclosure is not a concession. The deterrence evidence indicates that announced monitoring does most of its work before the attempt, so telling candidates strengthens the system's real function while giving the innocent the notice the Standards already say they are owed.
  • Track your flag-confirmation rate. The share of flags your own review sustains is your positive predictive value measured on your own population, the number no vendor can print. If nearly every flag confirms, the threshold is too timid and cheats are walking past it; if almost none do, the queue is noise and innocent candidates are absorbing it. Either result is actionable, and both are invisible to an organization that never reviews.
  • Put the false-accusation cost in the ledger. The cost of cheating is already on the risk register. Enter the other column beside it: candidacies wrongly ended, claims exposed, reviews staffed, and reputation spent among the unflagged. A threshold is a business decision about which column grows, and it should be set by someone who has looked at both.

A fifth discipline sits underneath the four: keep the meaning of a flag stable across the organization. The moment a recruiter treats "flagged" as a synonym for "cheated", every computed number in this article starts operating unsupervised, at whatever error rate the settings imply. Language is a control surface here. A flag is a question addressed to a process, and institutions that phrase it that way in their tooling, their templates, and their training are the ones whose answers hold up. The same stability protects the program itself: a review run consistently is the difference between an integrity function the business trusts and one it quietly routes around.

The two letters that opened this article were priced by everything in between. The first is cheap until it is challenged, and then it has nothing behind it: no evidence the candidate can see, no review a regulator can inspect, no reply anyone recorded. The second costs a review process, and buys the only integrity outcome worth having, cheats confirmed on evidence, the innocent cleared without a scar, and a decision that reads the same on a dashboard, in a deposition, and in the candidate's memory of the company.

Where 5Profiler stands

On 5Profiler, a flag can be traced from the number on the screen back to the moment that raised it. Every integrity signal carries its captured evidence to a human reviewer, and nothing is invalidated by software alone. Integrity scoring is transparent arithmetic rather than a black box. Detection is layered where fraud concentrates, and review begins where detection ends, which is the division of labor this article has argued for: software to raise questions, evidence to inform them, people to answer them. The two letters that opened this article are, in the end, a choice about that division, and the platform is built for organizations that intend to send the second.

Read the science behind the platform · See it on your roles

References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
  2. Cizek, G. J. (1999). Cheating on tests: How to do it, detect it, and prevent it. Mahwah, NJ: Erlbaum.
  3. Gilliland, S. W. (1993). The perceived fairness of selection systems: An organizational justice perspective. Academy of Management Review, 18(4), 694–734.
  4. Hausknecht, J. P., Day, D. V., & Thomas, S. C. (2004). Applicant reactions to selection procedures: An updated model and meta-analysis. Personnel Psychology, 57(3), 639–683.
  5. International Test Commission. (2014). International guidelines on the security of tests, examinations, and other assessments. International Test Commission.
  6. Wollack, J. A., & Fremer, J. J. (Eds.). (2013). Handbook of test security. New York: Routledge.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.