Buried in the contract that closed your assessment vendor deal is a security section, and it is almost certainly a list of features: lockdown settings, camera checks, flag reports. Nouns, bought once, evaluated at signature. What the list does not say is who acts when one of the nouns fires. Test security, as the testing profession defines it, is not a property of the software; it is the program run around the software: risks assessed before incidents arrive, detection someone actually watches, responses defined before they are needed, and the whole arrangement reviewed on a schedule. The distance between the list and the program is about a page of writing, and most contracts never write the page.
Run the tape forward to the day the list is tested. Word arrives from a proctor, a candidate, or the vendor's own monitoring: a form may be circulating, a batch of sessions shares an anomaly, a score already acted on no longer resembles the person it was attached to. In most organizations, the hours that follow are spent not on the incident but on archaeology of the org chart. The vendor assumed the employer would decide; the employer assumed the vendor would detect; legal wants to know who approved the monitoring in the first place; and a candidate is waiting on an answer that nobody owns. No control was missing. A sentence was.
This article is about that sentence and the program it belongs to. The controls themselves, identity verification, session monitoring, content defense, statistical review, are mapped in the companion article on remote assessment fraud and are not re-argued here. What concerns us is the governance wrapped around them: who assesses the risk, who reads what the controls produce, who owns each step of an incident, and what forces the loop to close. Professional guidance for all of this exists, has existed for a long time, and contains no mysteries. Most buyers have simply never asked to see it.
Test security is a program the profession already wrote down
The central document is short and public. The International Test Commission's guidelines on the security of tests, examinations, and other assessments (International Test Commission, 2014) frame security not as a set of purchasable features but as an ongoing program: assess the risks to your particular operation, prevent what can be prevented, detect what cannot, respond under procedures defined in advance, and review the whole arrangement on a schedule, with documented procedures and defined responsibilities at every step. The register of the guidance is not alarm but housekeeping. It reads the way a fire code reads, not because fires are common but because the time to decide who holds the extinguisher is before the smoke.
Around it sits a settled professional literature. Wollack and Fremer's (2013) Handbook of Test Security is the standard reference on securing testing programs, and its table of contents is the tell: threat assessment, prevention, detection, response planning, and documentation, treated as the ordinary chapters of running assessments rather than as emergency appendices. Cizek (1999) supplied the practitioner's ground truth years earlier: item exposure compounds over time, which makes content monitoring and refresh continuing obligations of a program rather than one-time purchases.
The profession's baseline document closes the frame. The Standards for educational and psychological testing (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014), the joint reference against which serious assessment practice is judged, treat security procedures as part of proper test use itself. An organization that interprets scores has obligations to protect the conditions that make the scores interpretable, and test takers hold rights to fair procedures inside whatever security regime surrounds them. In this framing, security sits inside validity rather than beside it: part of the conditions a score needs in order to mean anything.
Figure 1 draws the cycle these documents share: assess, prevent, detect, respond, review, and back to assess. The loop shape is the argument. Each station feeds the next, the last feeds the first, and none of the five is a thing you can finish buying.
What is striking about this literature is its circulation, not its content. It is more than two decades mature, freely available, and almost never referenced in procurement. Buyers who would not sign a payroll provider without asking about reconciliation routinely sign assessment contracts without asking to see the vendor's security documentation, and vendors are rarely pressed to volunteer it. The result is a market in which the assessment security program usually exists only as fragments, each side assuming the other holds the missing pieces.
A risk assessment starts with what is worth stealing
The cycle's first station is the one most organizations skip, because it feels as if it must already have been done by somebody. A security risk assessment for a hiring funnel answers three questions in order. The first is what the funnel holds that is worth stealing or faking: item content whose value survives one sitting, fixed forms that repeat across candidates, identities and credentials, score reports that open doors. The second is where scores actually drive decisions: at which stages, against which cut points, with what possibility of appeal. The third is which roles carry stakes beyond the hire itself: privileged access, regulated duties, safety-critical work, the positions where a faked score becomes an operational exposure rather than a hiring mistake.
The answers are local, which is why they cannot be inherited from a vendor's brochure. A retailer screening seasonal staff and a bank selecting traders may license the same instrument and face different threat pictures entirely; what is worth attacking is a fact about the decision, not about the test. The assessment also fixes the choice of administration mode. The profession codified a taxonomy of supervision levels for internet-delivered testing (International Test Commission, 2006), and matching mode to stakes, the subject of the companion article on unproctored testing and verification, is the natural output of the exercise described here: the risk assessment tells you which decisions can tolerate unsupervised delivery and which cannot.
Done properly, the exercise stays small. It is a working session and a memo, not a consultancy engagement: an inventory of what the funnel holds, a map of where its scores bind, a ranking of which failures would matter most. Its purpose is targeting. Prevention and detection consume attention, and the risk assessment decides where that attention sits before an incident announces where it should have been sitting all along.
Governance adds coverage, not controls
Prevention and detection are where the buying happens, and they are deliberately not this article's subject. The layered architecture, identity, session, content, and statistical review, belongs to the fraud article, and how those layers are engineered into a delivery system is a platform security question. What governance adds on top of the controls is three disciplines that no feature list can contain, because each one is a habit rather than a capability.
The first is coverage review. Controls cluster where products point them, which is historically the live sitting; the funnel stages before and after it, registration, identity claims, score reporting, record keeping, accumulate quietly and often carry no detection at all. A coverage review walks the funnel end to end and records, for each stage, what would notice an attack there today. What it produces is a short list of stages where the answer is nothing, and that list, not the feature sheet, should set the next procedure or the next purchase.
The second is a named reader. Detection that nobody reads is archival, not operational: a monitoring feed that no one is assigned to review is a record of incidents, not a defense against them. The governance question is not whether telemetry exists but whose calendar it appears on, at what interval, with what obligation to escalate. A standing half hour with the flag queue and the anomaly summary converts a product feature into an operating practice, and it is the cheapest conversion in this article.
The third is content discipline. Cizek's (1999) observation has not aged: exposure compounds, because every sitting teaches the candidate pool a little more about the test, and content that was secure at launch decays toward public knowledge on a schedule set by volume. Refresh and rotation are therefore maintenance obligations of the program, on a calendar, with an owner, exactly as certificate renewal is a maintenance obligation of an IT estate. A buyer who asks a vendor how content is retired and replenished, and on whose schedule, learns more about that vendor's security maturity than any feature demonstration will show.
Together the three disciplines answer the question a feature list cannot: not what the organization bought but whether anyone runs it. Coverage review shows where the controls end, the named reader shows whether detection is alive, and content discipline shows whether the asset being protected still holds its value. An audit that established only those three facts would already have earned its fee.
The incident half of the contract is usually unwritten
Everything so far can be run well and still fail on the day it matters, because response is the one station of the cycle that cannot be improvised without cost. The incidents at issue here are program-level: a form confirmed to be circulating, a batch of sessions sharing an anomaly, a systemic irregularity in one region or one recruiter's pipeline. The individual integrity flag, one candidate and one review, has its own process and its own error economics, covered in the companion article on integrity flags and false positives; this section is about events that touch many scores at once. For those, the practical question is never what to do, which the literature settles, but who does it, which only the contract can settle.
Time behaves differently inside an incident, and that is the argument for ink. Containment is measured in hours: a circulating form loses nothing by being withdrawn tonight, and preserved session records lose everything by being overwritten next week. Investigation is measured in days, and candidate decisions in whatever the offer calendar allows. An ownership split negotiated after the clock starts consumes exactly the hours containment needed, which is why the guidance insists that procedures and responsibilities be written before they are used (International Test Commission, 2014).
Figure 2 lays out the split that has to exist somewhere and is better written than assumed. Containment sits naturally with the vendor: the compromised form, the affected sessions, and the delivery infrastructure are in the vendor's hands, and the first hours belong to whoever can pull the levers. Investigation is joint, because the vendor holds the telemetry and the employer holds the context: which roles, which stakes, which decisions have already acted on the affected scores.
Consequences for a candidate belong to the employer, always. The moment a score touches an employment outcome, the decision is an employment decision, and no vendor should be making it. Communication follows the same logic, because candidates are the employer's relationship, informed by what the vendor's investigation established. Re-scoring and re-running are joint mechanics, the vendor supplying the instrument and the employer supplying the priorities and the schedule. Documentation and review are joint by definition, since the record has to serve both sides when the incident is examined later.
The grid is a default, not a doctrine, and moving its cells on purpose is the entire value of writing it down. What does not vary is the candidate's position inside the process. The Standards are direct on this point: test takers hold rights to fair procedures, and an incident does not suspend them (American Educational Research Association et al., 2014). A candidate whose score is caught in a compromised batch is owed what a flagged individual is owed: notice, an explanation, a path to re-demonstrate what the tainted sitting was supposed to measure, and a decision made by a person. Writing the ownership grid is partly a service to that candidate, because an owned step is a step that happens.
Review is what separates a program from a purchase
The last station closes the loop, and it is both the cheapest to run and the easiest to skip. Review means three standing habits. The first is scheduled: at a fixed interval, the program's owners look at what detection produced, what the flag queue confirmed and what it cleared, which is the confirmation-rate discipline the companion piece on integrity flags develops, what incidents occurred, and whether the coverage map's blank regions have moved. The interval matters less than the fixture. A review that lives on a calendar survives personnel changes; one that lives in intentions does not.
The second habit is the postmortem that changes something. An incident that ends with a closed ticket has been absorbed; an incident that ends with a modified control, a re-drawn coverage map, or a re-negotiated ownership cell has been used. The test of a postmortem is a diff — name the thing that is different now. A program that cannot name it is running response as customer service, and its next incident will look exactly like its last.
The third habit is triggered rather than scheduled: re-assess when the stakes change. The commonest version is silent promotion. An assessment adopted for early screening, where the posture the field judged defensible for unproctored delivery was explicitly conditioned on the modesty of the decision (Tippins et al., 2006), drifts into final-round or sole-source use because it is already there and already paid for. Nothing about the software changed; everything about the risk did. In the framing of the security guidelines and the Standards alike, a screening tool promoted to a deciding tool is a new security question, and the review habit is what forces that question to be asked out loud rather than discovered inside an incident.
The climb from feature to program is shorter than it looks
Locating your organization on this subject requires no audit; Figure 3 names four postures, and one of them will read as a description. At the first, a control has been bought and nobody owns it: proctoring is on, flags accumulate, and the security section of the contract has not been read since signature. At the second, modes are matched to stakes: someone has made the mode-to-stakes decision deliberately, and supervision varies with what each score is for. At the third, the program operates: flags are reviewed by named people, incidents have owners, telemetry has a reader. At the fourth, the program is reviewed: the calendar exists, postmortems produce diffs, and a change in stakes triggers a fresh look.
Two things about the four postures are worth saying plainly. Each step keeps everything from the step below, so nothing is discarded on the way up. And no step is bought. The move from the first posture to the fourth involves no new procurement at all: it is a page of ownership, a calendar of review, and a handful of named readers, applied to controls the organization already holds. That is why the right frame for this subject is maturity rather than exposure. An organization at the first step is not negligent; it is normal, and one working session away from the second.
Four lines for the contract, one date for the calendar
For organizations buying or operating assessment at scale, the program described above comes down to four moves, none of which requires a budget line.
- Write the ownership grid into the contract. Take Figure 2 as the draft, move cells where your situation argues for it, and make the result an exhibit rather than an email thread. The negotiation itself is diagnostic: a vendor who cannot say who contains a compromised form has told you which posture they occupy.
- Ask for the vendor's security documentation and incident procedure. The guidelines expect a testing program to hold documented procedures and defined responsibilities (International Test Commission, 2014), so the request is not exotic, and a vendor running a real program has the documents to hand. Any serious evaluation of an enterprise assessment platform should include it, and a test security audit, whether yours or a third party's, starts from those documents rather than from the feature list.
- Schedule the review before the first incident. Put the standing review on the calendar at signature, name its attendees from both sides, and fold its outputs into a written exam security policy that names owners. Policy that exists only as vendor settings is configuration, not governance, and it evaporates with the administrator who chose the settings.
- Treat security posture as part of validity. The Standards place security procedures inside proper test use because interpretation depends on them (American Educational Research Association et al., 2014); a score from a compromised process does not mean what the manual says it means. Procurement already asks for validity evidence. The same conversation should ask how the conditions for validity are protected in operation.
The distance between the lock and the program has been the subject throughout, and it is worth ending on its size. The literature is written, the controls are bought, and the people exist. What is missing, in most hiring funnels, is a page — who watches, who investigates, who decides, who tells the candidate — and a date on which the page is reread. Organizations write that page for payments, for privacy, and for the physical keys to the building. The scores that decide who joins the organization deserve the same ordinary diligence, and the profession that builds the tests wrote the instructions long ago.
Where 5Profiler stands
Bring your security team to the evaluation. 5Profiler runs test security the way this article describes, as a documented program rather than a bundle of settings: sessions produce captured records, every incident path has an owner written down before it is needed, and the resulting audit trail is one a security or compliance function can inspect line by line. The ownership grid in Figure 2 maps onto the platform's operating model directly, vendor-side containment and telemetry on one side, employer-side decisions on the other, with the documentation both sides need to hold their cells. Ask for the incident procedure during the demo — the answer should be a document, not a reassurance.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
- Cizek, G. J. (1999). Cheating on tests: How to do it, detect it, and prevent it. Mahwah, NJ: Erlbaum.
- International Test Commission. (2006). International guidelines on computer-based and internet-delivered testing. International Journal of Testing, 6(2), 143–171.
- International Test Commission. (2014). International guidelines on the security of tests, examinations, and other assessments. International Test Commission.
- Tippins, N. T., Beaty, J., Drasgow, F., Gibson, W. M., Pearlman, K., Segall, D. O., & Shepherd, W. (2006). Unproctored internet testing in employment settings. Personnel Psychology, 59(1), 189–225.
- Wollack, J. A., & Fremer, J. J. (Eds.). (2013). Handbook of test security. New York: Routledge.