The most valuable event in online proctoring appears in no log, because it is an event that does not occur: the attempt a candidate decides not to make once they believe someone is watching. What the vendor demo shows is the other job, detection, which catches the second face in the frame, the glance that keeps returning to the same spot, the voice that should not be in the room. A proctor does both jobs at once, and nearly the whole argument about monitoring, for it and against it, is an argument about the one that can be filmed. Most of what monitoring earns, this article will contend, comes from the one that cannot.
The distinction matters because each framing dictates a deployment. An organization that buys monitoring as a detector specifies it like a camera system: maximum capture, minimum notice, success measured in flags produced. Deployed that way, the technology collects the worst of both available worlds. Candidates experience an ambush, with everything that does to how they regard the employer, and the flags land on a team that never budgeted for reviewing them, where each unreviewed accusation sits as a liability quietly accruing interest.
Read as a deterrence instrument instead, the same technology asks for the opposite deployment: announced rather than covert, proportionate to stakes rather than maximal, and judged by the attempts that never happen rather than the ones it catches. The case for that reading comes from two directions. Criminology has spent decades asking what actually stops people from offending, and its answer transfers to the testing room as an argument worth taking seriously. And the one experiment that has tested monitoring directly in assessment found the pattern that argument predicts. The companion review of remote assessment fraud maps the threat and the layered controls that answer it, and leaves the behavioral science of monitoring to this article; what follows is that account.
The deterrence literature converges on certainty, not severity
Criminology has been asking what actually stops people from offending for the better part of a century, much of it inside the rational-choice frame Becker (1968) supplied — expected cost equals probability of being caught times consequence — which isolates the only two levers any integrity regime can pull: make apprehension more likely, or make its aftermath worse. Nagin (2013), in a synthesis of deterrence research published in Crime and Justice, states the field's conclusion plainly: the evidence that deterrence works is far stronger for the certainty of apprehension than for the severity of punishment. People are stopped by the belief that they will be seen and caught; hardening the penalty once certainty is established adds little. The corollary that carries this article follows directly: visible, credible monitoring deters, because visibility is the input to perceived certainty.
The certainty finding has an uncomfortable implication for assessment policy, because severity is the lever organizations reach for first. Severity is cheap to escalate on paper: a sternly worded integrity policy, automatic disqualification, the threat of a note that follows the candidate. A regime of harsh consequences attached to invisible monitoring is precisely the configuration the deterrence literature expects to fail, because a candidate who does not expect to be seen discounts the whole threat, however fearsome the consequence attached to it. Testing programs tend to write severity into policy documents and leave certainty to procurement, and the literature says that ordering is backwards.
The transfer to assessment is easy to state. A candidate weighing a proxy sitter, a second device, or a helper beyond the camera frame is running a recognizable version of that calculation, and an announced monitoring regime raises the perceived probability of being caught at the precise moment the decision is being made, which is before the sitting begins. Figure 1 places the two mechanisms on the timeline of a single attempt. Deterrence operates in the decide phase, where its successes are invisible by definition; detection operates during and after the sitting, on the residue of attempts the deterrent failed to prevent.
That transfer deserves to be labeled for what it is. Deterrence research studies crime and punishment (policing, sanctions, offenders with far more at stake than a test score), and no part of it has measured an assessment platform; carrying its conclusion into a testing room is an argument by analogy, not a measurement. The analogy is unusually well-shaped, because assessment misconduct is planned rather than impulsive, the would-be offender is weighing costs at a defined moment, and monitoring concentrates on exactly that moment, conditions under which the certainty mechanism should operate at its best. A well-shaped analogy is still an argument. What testing itself has measured comes next, and it is one study.
Online proctoring's direct evidence is one experiment
Karim, Kaminsky, and Behrend (2014) ran that study, published in the Journal of Business and Psychology and described by its own authors as an exploratory experimental study of remotely proctored testing. Three of its findings organize everything a buyer needs to weigh. Test-takers who were remotely monitored cheated less than those who were not. Their test performance was not meaningfully harmed, which speaks to the standing fear that being watched degrades the scores of the innocent along with the behavior of the guilty. And the monitored group reported greater privacy concerns, a cost this article returns to below rather than waves away.
The label on that evidence deserves as much attention as the evidence. This is one exploratory experiment, not a literature, and an entire product category leans on it whenever a deterrence claim appears in a proctoring proposal. Repeating the qualifier is not a way of dismissing the study; it is what taking a study seriously looks like. It cannot, alone, establish how large the effect is, how durable it is, or how it varies across stakes and populations, and no quantity of vendor confidence substitutes for the replications that have not yet been run.
What the experiment can do is converge. A deep literature about a different setting predicts that watched people attempt less misconduct, and the one direct test in the right setting found exactly that, in participants who had never read a criminology journal. Two independent lines pointing the same way justify designing as if announced monitoring deters, provided the design does not require the evidence to be stronger than it is. That standard, build on the convergence without overselling it, governs everything that follows.
A deterrent nobody knows about is only a detector
The psychology of exam monitoring is perceptual before it is technical, and this is the property that discussions of proctoring effectiveness most often miss. A candidate is deterred by the monitoring they believe is present, reviewed, and consequential, not by the monitoring that exists. However sophisticated the stack, a watcher the candidate does not know about contributes nothing to perceived certainty, and perceived certainty is where the deterrent lives. From the certainty finding, three working conditions follow, drawn as Figure 2.
Visibility means the monitoring is announced: candidates are told that observation is part of the process, before the process begins. Credibility means the announcement is believed, and belief has a maintenance cost, because applicant communities compare notes, and a monitoring regime whose flags demonstrably go nowhere is re-priced by the next cohort as scenery. Sustaining credibility therefore depends on the review capacity behind the flags, the machinery examined in the companion article on integrity flags and false positives. Immediacy means the information arrives at the decision moment: a plain-language notice at scheduling and again at the start of the sitting, not a clause inside a terms-of-service document no candidate has read.
A notice that deters is concrete. It says what is captured, who reviews it, how long the record is kept, and what happens when something is flagged, in the order a candidate would ask the questions. Vagueness is often defended as flexibility, but a vague warning fails both audiences at once: too thin to move perceived certainty for the candidate weighing an attempt, and too thin to read as fair process for the candidate who was never going to make one.
The fourth chip in Figure 2 sits apart deliberately. Covert monitoring, undisclosed by design, cannot move perceived certainty, so it forfeits the deterrent entirely and keeps only the detector; every attempt it will ever catch is an attempt it did nothing to prevent. Its yield is compromised at the moment of capture, because a flag obtained without notice records a violation of rules the candidate was never shown. And covert regimes tend not to stay covert; when they surface, in forums or in disputes, they surface as ambushes, at which point the organization has paid the full candidate-experience price of being watched secretly and banked none of the prevention.
Announcement also does quiet work on the majority who were never going to cheat. An assessment is a contest for rank, and every candidate in it has a stake in the contest being clean; telling candidates the sitting is observed tells them their competitors are observed too. For the honest majority, monitoring that is plainly announced and visibly proportionate reads less as accusation than as assurance: the effort they are about to spend cannot be quietly outbid by a purchase. A deterrence regime protects the people it never has to deter, which is a sentence no detection regime can say about itself.
Proportionality settles the candidate-experience bill
The cost side of announced monitoring is real, and the same experiment that documented the deterrent documented the price: monitored test-takers reported greater privacy concern than unmonitored ones (Karim, Kaminsky, & Behrend, 2014). Dismissing that as squeamishness misreads how selection works, because reactions travel: in the meta-analytic summary of applicant-reactions research by Hausknecht, Day, and Thomas (2004), the impression a selection procedure leaves on candidates carries over into what they make of the employer behind it and into their intentions toward it afterward. A monitoring regime is part of the procedure leaving that impression, and it is often the part candidates remember.
The fairness literature explains what separates monitoring candidates accept from monitoring that reads as an insult. Gilliland (1993) identified procedural-justice rules (among them consistency of administration, transparency, and the opportunity to respond) that predict how a selection process is received, and announced, evenly applied monitoring with a channel for explanation satisfies those rules in a way covert or arbitrary monitoring cannot. The candidate experience of remote proctoring, in other words, is not a fixed toll; it is largely a function of deployment choices the buying organization controls.
The choice that controls most is level. Proportionality, monitoring at the level the stakes justify rather than the level the technology allows, resolves the tension between deterrence and experience by refusing to treat it as all-or-nothing. Figure 3 draws the ladder. A practice sitting carries no decision and warrants no observation. A high-volume screen justifies automated presence checks, enough to make the announcement true. A shortlist or final stage, where scores carry offers, warrants a live supervised sitting. A high-stakes certification, where the credential outlives the sitting, can justify a fully managed environment. The candidate-experience cost climbs the same ladder, which is exactly why each step should be spent only where the decision requires it.
The level the technology allows is a moving target, which is what makes proportionality a standing discipline rather than a one-time setting. Every release season adds capture the previous one lacked, and each new capability arrives with an implicit argument that using it everywhere is safer than using it somewhere. The question a buyer should put to each addition is not whether it strengthens detection, which it usually does, but which rung of the stakes ladder needs it, because a capability adopted without a rung assigned migrates toward everywhere by default.
Proportionality is also worth announcing in its own right. A candidate asked to sit a fully supervised final round accepts it more readily when the earlier screen was visibly lighter, because the escalation itself communicates that observation tracks the decision rather than the organization's appetite. The ladder, in other words, is legible from the inside. Candidates who can see stakes and scrutiny rising together are being shown the consistency and transparency the fairness rules reward, one sitting at a time.
Two boundaries keep the ladder in its lane. Which administration mode a sitting runs in, and where supervised verification of scores belongs in a two-stage design, is territory mapped in the companion article on unproctored testing and verification; the ladder here sets how closely any given sitting is watched once that design is in place. And the ladder is deliberately generic: the monitoring levels a given delivery platform exposes are a product question (5Profiler's are described on the features page), while the pairing of level to stakes is a policy the buying organization must own and be able to recite.
Detection's product is evidence, not verdicts
Pricing deterrence first clarifies what detection is for. Some minority of candidates will attempt misconduct against any announced regime, and for them detection does the one job deterrence cannot: it captures a record. Cizek (1999), writing on cheating on tests before webcams existed, put prevention and deterrence first in test security and treated detection statistics as triggers for investigation rather than verdicts, and nothing about software has amended that ordering. A flag is the opening of a case, not its closing: captured evidence, a human reviewer, a proportionate response, and a conclusion that can be shown to the person it concerns. Review is also where the deterrent gets its credibility serviced, because every case examined and closed on evidence is what makes the next cohort's announcement believable; the reviewer is doing preventive work even when the file in front of them turns out to be a false alarm.
What a flag is statistically worth, why the answer is worse than most buyers assume, and what a review pipeline that survives challenge looks like are the subject of the companion article on integrity flags and false positives. The arithmetic belongs there; the design consequence belongs here. A program that announces its monitoring, deters most attempts, and reviews what remains produces a short queue of well-evidenced cases, which is the only kind of queue an organization can defend. And the point of the whole apparatus is narrower than the apparatus suggests: the instruments being protected, whose validity case is laid out on the science page, only measure anything when the person measured is the candidate.
Design for the attempt that never happens
The deterrence framing converts into four operating rules for anyone running online proctoring on a hiring funnel.
- Announce monitoring in plain language, before every sitting. Say what is observed, why, what happens when something is flagged, and who reviews it. The announcement is the deterrent's delivery mechanism, not a courtesy attached to it, and burying it in fine print discards the deterrent while retaining the intrusion.
- Match the monitoring level to the stakes, and put the mapping in the assessment policy. A recorded pairing of assessment stage to monitoring level, in the spirit of Figure 3, turns proportionality from a sentiment into a commitment that procurement, counsel, and candidates can all read and hold the program to.
- Never run covert monitoring on a hiring funnel. Undisclosed observation prevents nothing, charges the full privacy price, and generates flags obtained under rules candidates never saw. Whatever covert capture is worth elsewhere in security practice, on a hiring funnel it is a purchase of liabilities.
- Audit the deterrent, not just the detector. Flag volume measures activity. If announced monitoring is doing its work, successive cohorts should attempt less, and the flag rate should drift down as the regime becomes known and believed.
The last rule needs its own picture, because every dashboard argues against it. Figure 4 contrasts what monitoring software reports, a volume of flags that reads like productivity, with what a working deterrent actually leaves in the data: a flag rate that falls across successive cohorts as announcement and review become common knowledge. One caution keeps that metric from becoming a target: a falling rate is also what a decaying detector or a quietly loosened threshold produces, so the trend should be read by whoever owns the regime's design, with detector settings held steady across the cohorts being compared, and with outcomes from supervised verification sittings as the cross-check. But the direction of the goal should not be controversial. A proctoring program whose flag rate never declines, cohort after cohort, is paying for a deterrent and receiving only a detector.
Which returns to the two jobs. The demo will keep selling the catch, because the catch can be filmed, and buyers will keep comparing online proctoring by the sharpness of its detectors, because sharpness fits in a comparison table. The larger value does not photograph: candidates told the truth about being observed, at a level the stakes can justify, deciding one at a time that the attempt is not worth making. A monitoring program succeeds twice. Once, visibly, in the small queue of flags its reviewers can defend; and once, invisibly, in the queue that never formed.
Where 5Profiler stands
Monitoring on 5Profiler is matched to stakes and announced to candidates. Assessment owners set the monitoring level per stage, from unobserved practice sittings to fully supervised high-stakes sessions, so candidates carry only the observation the decision in front of them justifies. What is monitored is disclosed before the sitting begins, in plain language rather than a terms-of-service clause, because an announced watcher is the one that deters. Proportionality here is a configuration surface, not a promise: the stake-to-level pairing this article recommends is something an assessment owner sets, sees, and can show. The success metric worth watching is the article's closing image: the queue that never forms.
References
- Becker, G. S. (1968). Crime and punishment: An economic approach. Journal of Political Economy, 76(2), 169–217.
- Cizek, G. J. (1999). Cheating on tests: How to do it, detect it, and prevent it. Mahwah, NJ: Erlbaum.
- Gilliland, S. W. (1993). The perceived fairness of selection systems: An organizational justice perspective. Academy of Management Review, 18(4), 694–734.
- Hausknecht, J. P., Day, D. V., & Thomas, S. C. (2004). Applicant reactions to selection procedures: An updated model and meta-analysis. Personnel Psychology, 57(3), 639–683.
- Karim, M. N., Kaminsky, S. E., & Behrend, T. S. (2014). Cheating, reactions, and performance in remotely proctored testing: An exploratory experimental study. Journal of Business and Psychology, 29(4), 555–572.
- Nagin, D. S. (2013). Deterrence in the twenty-first century. Crime and Justice, 42(1), 199–263.