Unproctored testing began as a deliberate trade, struck knowingly and argued in public at the moment employment testing left the supervised testing center for the candidate's browser. Organizations bought reach, speed, and candidate convenience at a scale no proctored program could match, and they paid in certainty — certainty about who sat the test, under what conditions, and with whose help. The research on that trade supports a conclusion neither camp in the proctoring argument fully enjoys: the defensible design is not surveillance everywhere or trust everywhere but administration mode matched to stakes, with supervised verification spent precisely on the scores an organization is about to act on.
The difficulty is that most organizations never made the trade explicitly. They inherited a mode of administration from a vendor default or a procurement precedent, and they now run it for every assessment at every stage, from practice quiz to the final gate before an offer. The single-mode habit produces both of the failure patterns this article addresses. Some programs over-proctor, loading surveillance cost, candidate friction, and privacy exposure onto the top of a funnel where stakes are low and volume is high. Others under-verify, acting on unsupervised scores at exactly the point where a decision makes those scores worth corrupting.
What follows assembles the vocabulary the field built for this problem and the evidence for using it: the four modes of administration, the risks the founding debate predicted, what operational programs found when they measured score inflation, and the two-stage verification design that makes a cheap screen and an expensive certainty complements instead of rivals.
Four modes of administration, not a proctored/unproctored switch
The argument is usually staged as a binary, proctored versus unproctored, and the binary conceals most of the design space. The International Test Commission's guidelines on computer-based and internet-delivered testing replaced it with a taxonomy of online test administration modes, four of them, ordered by increasing control over identity and conditions (International Test Commission, 2006). Open mode is assessment available to anyone, anywhere: no supervision, and no confirmation of who is responding. Controlled mode restricts access to registered candidates, but the sitting itself remains unobserved. Supervised mode puts a proctor in the loop, in person or remote, to confirm identity and testing conditions. Managed mode adds control of the environment itself: dedicated testing facilities, known machines, standardized procedure.
Two properties of the taxonomy do most of its practical work. First, mode attaches to the sitting, not the instrument: the same assessment can run in open mode for one purpose and supervised mode for another, which is what makes administration a design variable rather than a fact about the product. Second, the four modes form a ladder, drawn in Figure 1. Each rung buys assurance the rung below cannot provide, and every purchase is paid for in something the hiring funnel values: reach narrows, friction rises, scheduling appears, cost per sitting climbs. The ladder converts an argument about principle into a decision about placement. "Should we proctor?" has no answer; "which rung does this decision require?" always does.
Most buyers have never been shown these four modes side by side. Platforms tend to present whichever mode they deliver as the natural way to test, and procurement tends to ask whether proctoring exists rather than which modes can be assigned to which stage of the funnel. The taxonomy is this article's spine because everything that follows happens on it: the founding debate was about which rungs are defensible for which purposes, the operational evidence is about what actually occurs at the unsupervised end, and verification testing is a scheduled climb, made by exactly the candidates whose scores are about to matter.
The field predicted the risks of unproctored testing at the moment of adoption
Unproctored internet testing received its trial in public. Tippins and colleagues (2006), in the focal article in Personnel Psychology that framed the debate, convened the arguments for and against the practice while it was spreading through operational use, and the risk catalogue they assembled has not aged. When nobody observes the sitting, candidate identity cannot be assured. Items exposed to unsupervised candidates leak into circulation, and the leakage compounds with every additional sitting. And assistance, from a helpful friend at the keyboard to whatever reference material sits within reach, becomes an option each candidate must individually decline.
What the panel did with the catalogue matters more than the catalogue. Its authors disagreed with one another in print, and the disagreement converged rather than dissolved: unproctored delivery can be defensible for screening, under conditions, and is not defensible as the basis for a final decision unless the scores that carry the decision are verified under supervision. The condition was not a truce between optimists and pessimists; it was a design both could sign, because it concedes the lower rungs' weakness on identity while spending supervision's cost only where that weakness becomes intolerable.
The risk list also has an internal order that repays attention. Identity and assistance are per-candidate risks: each sitting either was or was not the named person, helped or unhelped, and the harm stays attached to that score. Exposure is a program risk: one harvested form quietly degrades every future sitting, including the supervised ones, because a memorized answer key works in front of a proctor. And the founding debate was mostly imagining opportunistic assistance, a friend, a browser tab, a calculator where none was allowed. How the threat has changed since, including the market for outside help, is the subject of the companion article on remote assessment fraud linked below.
Item exposure deserves one more sentence, because it compounds where the other risks merely occur. A fixed form delivered unproctored is a publication event on a delay: every sitting is an opportunity for content to be harvested, and harvested content converts future sittings from measurement into recall. The structural answer, calibrated item banks with per-candidate adaptive sequences, belongs to the series' companion article on adaptive testing; what matters here is that exposure is a program-level risk rather than a per-candidate one, and that it accumulates fastest in the unsupervised modes.
Operational programs found modest inflation, not the feared collapse
The 2006 debate ran ahead of the data, but not by much, and the data that arrived came from working programs with real applicants and real stakes rather than from laboratory volunteers with nothing to gain. That provenance matters: the operational studies observed candidates who had every incentive the pessimists worried about, in the settings the design was built for. Nye, Do, Drasgow, and Fine (2008) examined an operational two-step program built on the panel's design, an unproctored screen followed by proctored verification, and found little evidence that score inflation was a meaningful problem in practice. Lievens and Burke (2011) then reported results from a large-scale operational testing program in the United Kingdom, cognitive ability delivered unproctored at applicant volume: score inflation at the unproctored stage was modest, and verification testing proved workable as an operating routine rather than a pilot exercise.
Arthur, Glaze, Villado, and Taylor (2010) approached the question from the comparison side, setting unproctored internet-based test scores against verified retests. On the cognitive measures, the mean differences were smaller than the debate had braced for. The same study found something less comfortable in the distribution: a minority of candidates showed score patterns consistent with cheating. Both findings are true at once, and the pairing, small average effects over a genuinely misbehaving minority, is the pattern a buyer of unproctored assessment needs to hold in mind.
One lane boundary should be marked before interpreting any of this. The findings above concern ability testing in unproctored settings; applicant score inflation on personality measures is a separate literature with separate mechanics, treated in the series' companion article on faking, and the two should not be averaged into a single anxiety.
Read as a panel (Figure 2), the three findings agree in direction, and the direction is reassuring: the unsupervised modes did not collapse under operational weight. But the reassurance is specifically an average's reassurance, and it has two sharp limits. Modest inflation is not no inflation; the phrase records a real effect that stayed small in the aggregate, not an absence. And selection does not consume aggregates: shortlists are drawn from the tail of the score distribution, where a small minority with inflated scores can be heavily overrepresented without disturbing the mean, a tail argument developed, with the computation, in the companion article on remote assessment fraud.
What the operational record supports is therefore narrower than either camp's talking points. It supports the unproctored screen, because the average distortion stayed modest in well-run programs. It declines to support the unproctored decision, because the misbehaving minority exists, concentrates where selection happens, and is invisible to the mean. Between those two sentences sits the two-stage design.
Verification testing spends certainty only where decisions are made
The design the founding panel described and the operational programs ran has two stages (Figure 3). Stage one is an unproctored screen, open or controlled mode, doing what the lower rungs do best: reaching every applicant, at applicant volume, on the candidate's own schedule, at low cost per sitting. Stage two is verification testing, a supervised re-assessment (supervised or managed mode) of the candidates whose stage-one scores the organization intends to act on. The screen produces a shortlist; the verification produces the evidence the decision will actually rest on.
The economics are the argument. Supervision is the expensive ingredient in any testing program, and its cost scales with the number of people supervised, while a funnel's shortlist is a small fraction of its pool. The two-stage design therefore buys certainty where it is scarce and consequential, at the decision, and declines to buy it where it is plentiful and cheap, at the screen. Candidate experience follows the same gradient. The many get a convenient sitting from home; the few asked to bear verification's friction are the ones with an offer in view, which is where friction is best tolerated and where care is most owed. Logistics follow it too: verification can run as a remote supervised session or fold into an interview day the shortlisted candidate was attending anyway, so the climb up the ladder need not add a separate trip.
The pre-internet default was single-stage supervision: everyone who tested, tested under someone's eye, which priced measurement out of the top of the funnel and left the early sorting to weaker filters. The two-stage design did more than cut cost. It moved real measurement earlier, to the full applicant pool, where those filters used to do the deciding, and that is the gain a blanket return to supervision would surrender. It is also why the answer to integrity anxiety about unproctored testing is a second stage, not a retreat from reach.
The step most often misread in this design is the discrepancy, a supervised score that lands materially below its unproctored predecessor. A discrepancy is a flag for structured review, not a verdict. Sittings under different modes differ in conditions, equipment, familiarity, and nerves, any of which can move a score without misconduct, and an unproctored score can also err in a candidate's favor for reasons well short of cheating. The defensible response is a review of the evidence under a defined procedure, with the possibility of an innocent explanation held genuinely open; why integrity flags mislead when they are treated as verdicts is the subject of the companion article on integrity flags and false positives.
Two further design details keep the second stage from undoing the first. Content should not repeat between stages: a verification sitting run on the same form measures memory of the screen and feeds the exposure problem, which is why per-candidate sequences drawn from item banks, per the adaptive-testing article above, are the standard pairing with two-stage programs. And the second stage needs standing rather than improvisation: verification retesting is among the controls the International Test Commission's test-security guidelines recommend for scores used in decisions (International Test Commission, 2014), and the standing program that keeps such controls owned, audited, and rehearsed is the subject of the series' companion article on assessment security governance.
Where unproctored assessment is enough, and where it never is
The matching logic runs in both directions, and the direction the integrity literature rarely bothers to argue is the one that saves money. For practice tests, self-assessment, and development feedback, open and controlled modes are the correct choice rather than a fallback: stakes are low, candidates have little incentive to cheat themselves, and every unit of supervision added is pure cost, in money, in scheduling, and in the privacy exposure of putting cameras into homes where nothing decision-relevant is happening. Over-proctoring at the top of the funnel is not rigor. It is a tax paid by the many, mostly in goodwill, to secure decisions that have not yet arisen. And the privacy line item is real: supervision at home means cameras in bedrooms and kitchens, identity documents held up to webcams, and recordings retained somewhere, collected from a population of whom the organization will hire almost nobody. Spending that imposition on a practice quiz buys nothing any decision needs.
Screening is the middle case, and the operational evidence above says how to run it: the lower modes are acceptable for screening precisely when verification stands behind them. The assurance is borrowed from stage two. Deleting the verification stage does not so much simplify the design as convert the screen, silently, into a decision instrument the founding debate declined to defend.
The never case follows: final-decision scores that nobody re-verifies. An offer extended on an unverified unproctored score is the arrangement the Tippins panel could not defend, and nothing in the operational literature since has rehabilitated it, because the studies that reassure about averages are silent about the specific candidate in front of you, and the decision is about that candidate. Remote testing integrity, in the end, is a property of a design rather than of a platform: the same open-mode sitting is perfectly sound as a practice instrument and indefensible as the sole basis for an offer.
Figure 4 renders the prescription as a grid. A program can locate every assessment it runs in one of its cells in an afternoon, and the exercise is worth doing before any vendor conversation, because the completed grid is the requirement sheet that conversation should start from.
Write the mode-to-stakes mapping down
The implications are unusually cheap to act on, because the deliverable is a page of decisions rather than a program of work. Four moves cover it.
- Map every assessment to a mode and a stake. For each instrument at each stage, record the administration mode it runs in and the decision its scores feed. Any row where an unverified open- or controlled-mode score drives a hiring decision is a finding, and the fix is a mode change or a verification stage, not a sterner policy memo.
- Verify what you act on. Supervised verification of shortlist scores is the control the founding debate converged on (Tippins et al., 2006), the control an operational program ran without meaningful inflation (Nye et al., 2008), and the control a large-scale program showed workable as routine (Lievens & Burke, 2011). It is the move that converts a screening signal into decision evidence.
- Tell candidates the design up front. Say at invitation that shortlisted scores are confirmed under supervision. An announced design is experienced as process rather than ambush, and announcement lets the second stage shape behavior at the first, deterrence ground covered in the companion article on the psychology of proctoring.
- Buy modes, not a mode. Ask vendors which administration modes their delivery platform supports at which stage, and whether one program can assign different modes to screening and to verification. A platform that answers with a single mode for everything is answering a different question than the one hiring asks.
One caution keeps the list from curdling into policy theater. The mapping is a living document only if someone owns it: modes drift as platforms update, stakes drift as a screening score quietly hardens into a cutoff, and a grid that was accurate at purchase can be fiction two hiring cycles later. Revisit it whenever an instrument, a stage, or a decision changes, and treat that review as the smallest unit of the standing discipline described in the governance article linked above.
The trade that created unproctored testing was a good trade, and the researchers who watched it happen said so while insisting on its terms: reach, speed, and convenience at the top of the funnel; certainty, purchased under supervision, for the shortlist. Organizations get into trouble not by making that trade but by forgetting they made it, running one mode for everything and discovering, one contested hire at a time, which corner of the grid they have been living in. The candidates you shortlist are the only ones whose scores need certainty. Spend the supervision there.
Where 5Profiler stands
Two-stage administration is native to 5Profiler, not an integration project. A single program can run its screening stage unproctored at full applicant volume, then schedule supervised verification for the shortlist, with the administration mode assigned per stage rather than per platform — the mode-to-stakes mapping this article recommends, written into the assessment itself. Adaptive testing keeps the verification sitting from being a rerun of the screen, so a confirmed score reflects fresh measurement rather than rehearsal, and a between-stage discrepancy opens a structured review rather than an automatic verdict. Buyers comparing platforms should ask the question this article ends on: not whether a vendor proctors, but which modes it supports, at which stage, for which decisions.
References
- Arthur, W., Glaze, R. M., Villado, A. J., & Taylor, J. E. (2010). The magnitude and extent of cheating and response distortion effects on unproctored internet-based tests of cognitive ability and personality. International Journal of Selection and Assessment, 18(1), 1–16.
- International Test Commission. (2006). International guidelines on computer-based and internet-delivered testing. International Journal of Testing, 6(2), 143–171.
- International Test Commission. (2014). International guidelines on the security of tests, examinations, and other assessments. International Test Commission.
- Lievens, F., & Burke, E. (2011). Dealing with the threats inherent in unproctored Internet testing of cognitive ability: Results from a large-scale operational test program. Journal of Occupational and Organizational Psychology, 84(4), 817–824.
- Nye, C. D., Do, B.-R., Drasgow, F., & Fine, S. (2008). Two-step testing in employee selection: Is score inflation a problem? International Journal of Selection and Assessment, 16(2), 112–120.
- Tippins, N. T., Beaty, J., Drasgow, F., Gibson, W. M., Pearlman, K., Segall, D. O., & Shepherd, W. (2006). Unproctored internet testing in employment settings. Personnel Psychology, 59(1), 189–225.