Sooner or later, every assessment program produces the meeting. A candidate has finished at 69 against a pass mark of 70. The hiring manager, who liked everything else in the file, wants the interview anyway, and someone has to defend the line. Not the test — the line. Why 70? Who decided that 69 means no, and on what evidence? Too often the answer is silence, because the mark arrived as a vendor default or a predecessor's round number, and no one in the room can reconstruct a reason for it.
The research on setting cut scores has an answer, and it is disarming: no method discovers the right pass mark, because no such number exists to be discovered. A cut score converts a continuous measurement into a verdict, and where you put the conversion is a judgment about which mistakes you would rather live with: admitting some people who will not reach acceptable proficiency, or turning away some who would. Every method in the standard-setting literature structures that judgment; the judgment itself stays (Kane, 1994).
That is better news than it sounds. Understood as policy expressed in measurement units, a pass mark becomes something an organization can govern: placed by a process that can be shown, tuned against consequences that can be counted, revised when the evidence moves. A mark set by open judgment can be defended. A mark set by nobody cannot.
Setting cut scores means choosing where the judgment sits
The field's working taxonomy sorts the methods for setting cut scores by exactly that question: where each one asks human judgment to be exercised (Cizek & Bunch, 2007). Three families cover nearly everything an employer will meet in practice, and Figure 1 sketches one of each.
Test-centered methods put the judgment inside the test. In the Angoff method, a panel of judges imagines a minimally competent performer, the weakest person who should still pass, and estimates, item by item, the probability that this person answers correctly (Angoff, 1971). Summing the estimates gives the score such a person would be expected to earn, and that expectation becomes the pass mark. The procedure's value lies in what it forces. Before anyone can rate an item, the panel must agree on what minimal competence means for the role, in writing, and every rating afterward is an auditable trace of that definition meeting the test's content. Serious Angoff studies train their judges, let them revise after seeing how items behave with real examinees, and report how far the panel agreed, because the mark inherits whatever care the ratings received (Cizek & Bunch, 2007).
Examinee-centered methods move the judgment from items to people. In the contrasting-groups design, the organization identifies people it independently knows to be masters and non-masters of the work, administers the test to both, and plots the score distributions of the groups. The distributions overlap; the passing score goes where the curves cross, the region where a score stops being more typical of one group than of the other (Livingston & Zieky, 1982). A variant, the borderline-group method, asks judges to nominate people whose work sits at the boundary itself and centers the mark on how that group actually scores. Livingston and Zieky's manual for the Educational Testing Service is plain about what these designs consume: the mark is only as good as the classifications behind it, and the judgment has migrated into deciding who counts as a master.
Choosing among the families is itself part of the judgment, and the sensible basis is what the organization can actually supply. An Angoff panel needs subject-matter experts who know the role well enough to hold a defensible image of minimal competence, and a test whose items they can read against that image. A contrasting-groups study needs something rarer: a population whose mastery is known from evidence independent of the test, which usually means incumbents with credible performance records. Where neither exists, the method degrades into ceremony, judges rating items they cannot picture or groups classified by reputation. Matching the method to the evidence on hand is an early defensibility decision in its own right, and the standard guides treat it as one (Cizek & Bunch, 2007; Livingston & Zieky, 1982).
The third family hands the judgment to the measurement model. Score banding treats scores as tied whenever they fall within a band whose width comes from the standard error of the difference, the uncertainty that attaches to the gap between two scores rather than to either score alone (Cascio, Outtz, Zedeck, & Goldstein, 1991). Banding gets a section of its own below, because it answers a different question: the other families place the line, and banding decides how thick the line can claim to be.
Kane (1994), writing in the Review of Educational Research, drew the conclusion that organizes the whole taxonomy: a passing standard is a judgmental construction, and validating one means showing that the process was reasonable, that its pieces cohere, and that the documentation exists, down to evidence about the standard's consequences. There is no convergence test to run, because there is no hidden value to converge on. Two competent panels can land on different marks without either being wrong. What separates a defensible mark from an indefensible one is whether anyone can show how it was made.
Raising the mark shrinks the pass list and strengthens it
Wherever the judgment sits, the placement it produces has consequences that are computable in advance. The machinery is the standard bivariate-normal model behind the Taylor-Russell tables (Taylor & Russell, 1939), the framework the companion briefing on the cost of a bad hire uses to express validity as mis-hire risk. Here it answers the policy owner's question instead: what any given position of the mark implies about who passes.
Assume scores and later job proficiency are correlated at the level of the test's validity, taken at .50 for illustration. Each position of the mark then fixes a pass rate and an expected proficiency among those who pass, and Figure 2 traces both quantities as the mark slides from well below the pool average to well above it. Set the mark at the pool mean and half the pool passes, with passers expected to sit about .40 standard deviations above average on later proficiency. Push the mark one standard deviation higher and 16 in 100 pass, while the passers' expectation climbs to about .76 SD. The curves move against each other the whole way: whatever the line concedes to throughput it gives up in expected proficiency, and the reverse.
A mark can be wrong about a person in two directions. It can pass someone who will not reach acceptable proficiency, an unqualified passer, or fail someone who would have, a qualified failure. The right-hand panel of Figure 2 maps both regions against the joint distribution of scores and later performance: raise the mark and the first region shrinks while the second grows; lower it and the trade runs the other way; no placement empties both. Which error matters more is a fact about the role. Where most applicants would succeed, a high mark mostly manufactures qualified failures; where few would, a generous mark mostly admits unqualified passers. The governing quantities have names: the base rate, the share of applicants who would prove proficient if hired, and the selection ratio, the fraction of the pool the organization intends to appoint. Both differ role by role, which is one reason a single company-wide cutoff score is usually a category error in hiring. And because applicant pools shift, a mark that stands still while the pool moves quietly becomes a different policy without anyone deciding it should.
The errors also differ in how visible they are. An unqualified passer is hired, struggles, and becomes a story the organization tells about the test. A qualified failure disappears at the gate, succeeds somewhere else, and is never observed at all. Decision-makers who learn only from the visible error will push the mark upward over time, trading away candidates who would have served for protection against a failure they can see, and the ratchet has no justification in the measurement; it is an artifact of which mistakes leave records. A cut-score review that counts both regions of Figure 2, rather than only the one with a face attached, is correcting for that bias.
Regulation asks for the same reasoning. The Uniform Guidelines on Employee Selection Procedures, adopted jointly by the federal enforcement agencies, direct that cutoff scores normally be "reasonable and consistent with normal expectations of acceptable proficiency" within the workforce (Equal Employment Opportunity Commission et al., 1978). That is an instruction, and a deliberately loose one: connect the line to the job. The SIOP Principles translate the duty into professional practice: a cutoff score should rest on a job-related rationale, and setting one may legitimately weigh the labor market, the workforce the organization expects to need, and the consequences of the placement, among them its adverse impact, the arithmetic the briefing on the four-fifths rule owns (Society for Industrial and Organizational Psychology, 2018). A mark placed with Figure 2's trade in view and a written rationale behind it is what both documents are asking an employer to produce.
Banding concedes what the instrument cannot resolve
Banding begins from a concession the rest of scoring practice rarely makes out loud: the instrument's resolution is finite. Every score carries measurement error, and the gap between two scores is uncertain twice over, once for each score inside it. The standard error of the difference gives that compound uncertainty a width: the smallest gap the test can treat as evidence that two candidates genuinely differ. How to read that error term when a single report lands on a desk is the subject of the companion briefing on measurement error for decision-makers; score banding is what happens when a scoring policy takes the width seriously.
The mechanics are short (Cascio et al., 1991): start from the top score, extend one band-width down, and treat everyone inside as tied. Figure 3 shows the result on a ranked pool. Candidate A holds the top score; Candidate E trails by four points in the figure's illustrative pool; the band says the instrument cannot certify that A is stronger than E, so a policy that appoints A over E on that difference alone is ordering noise. The scores below the band are another matter: those gaps exceed what the error term allows to be called a tie, and the ranking there stands on measurement.
Banding solves one problem and immediately creates its successor. If the test cannot order the band, something else must, and that something is a policy choice too. Left unstated, within-band selection drifts toward whatever the file happens to evoke, which is precisely the discretionary judgment scoring exists to structure. The remedy is pre-commitment: fix, at the moment the band is defined, what decides among tied candidates (further job-related evidence, a work sample, a structured interview score) and apply it identically to every tie. And the width has to come from the error term: stretch the band beyond what the measurement model licenses and real signal is surrendered along with the noise (Cascio et al., 1991).
Banding fits best where its premise bites hardest: dense score regions where candidates stack within a point or two of one another, verdicts that are hard to reverse, pools that produce ties by the dozen. Campus seasons and enterprise hiring programs meet all of those conditions at once, and they are where a band, and the rule that governs its interior, do the most work.
The converse cases matter as much. A sparse senior pool with clear gaps between candidates gains little from a band, because the instrument can already support the orderings the decision needs. And where the assessment is one input among several read together, formal score banding may matter less than simply carrying the score's uncertainty into the room. Banding is a policy for verdicts that scores must make alone, at volume, near a line; used there, it keeps the instrument from testifying beyond its competence.
The line survives scrutiny as a record
Kane's criterion makes defensibility an evidentiary standard rather than a mathematical one, and most of the literature on setting cut scores converges here. The showing has ingredients, and the field is unusually concrete about them (Cizek & Bunch, 2007): who judged, and what standing they had to judge; the definition of minimal competence they worked from; what data they saw, including how the items or the known groups behaved; how disagreement among the judges was handled; who approved the result; and when it will be revisited. The professional testing standards make the duty explicit: the rationale and the process behind a cut score are to be documented (American Educational Research Association et al., 2014). The list is mundane, and that is its virtue: every entry exists so that a stranger can reconstruct the judgment.
The record also needs an owner, because documents do not revisit themselves. Someone accountable holds the mark: convenes the review, watches the cohort numbers, and decides when drift has crossed from noise into signal. The review triggers can be written in advance, like the mark itself: a scheduled date, a change in the role, a new form of the test, a shift in the applicant pool large enough to move the pass rate. An organization that can name the owner, the last review, and the next one has already answered most of any future challenge.
The record does not close at adoption. Kane's framework counts consequences as validity evidence, and consequences are observable: pass rates by cohort and by season; drift as the applicant pool changes; and the later performance of people who passed near the mark, the closest thing to a field test of the line an employer will ever run. A mark whose near-mark passers reach acceptable proficiency is corroborated by them. A mark whose near-mark passers struggle is telling you where to look, and an organization that never looks is running its standard on faith.
This is where the distinction that decides real disputes lives. A line placed by structured judgment is arbitrary in the strict sense: a different competent panel might have put it a few points away, and no measurement can adjudicate between them. It is not capricious. Capricious is the other condition: a line with no author, no definition of the person it was meant to pass, no memory of why here and why not elsewhere. Standard-setting methods do not exist to escape the first condition, which is inescapable. They exist to prevent the second, and the prevention is the process, written down. The account is quick to produce for anyone who kept it, and impossible for anyone who did not.
Treat the pass mark as standing policy
The working implications follow directly, and the first concerns inheritance: no default mark deserves to survive unexamined. A vendor's recommended passing score encodes assumptions about applicant pools, base rates, and error tolerances that are not necessarily yours. Restating the default in the role's terms is a small exercise with the organization's own data: what share of last season's applicants would have cleared it, how the people who cleared it narrowly went on to perform, what the minimally qualified person for this role looks like in the Angoff sense. A configurable assessment platform makes the restatement operational, one threshold per role, on the organization's own definition of competence, in place of one mark per product.
Second, the mark should be set with its error band in view. A score sits inside an interval, and a verdict that flips on a difference smaller than the instrument can resolve is a verdict the instrument never issued. Where candidates stack densely near the line, decide the within-band rule at the same meeting that fixes the band, while the tie is still an abstraction and not a person anyone has met. Fixed early, the width protects the process from its own sympathies; fixed late, it becomes a lever for whoever liked the file.
Third, write the rationale memo the day the cut is set, not the day it is challenged. Kane's coherence cannot be manufactured retroactively, and a justification reconstructed under pressure reads like what it is. The memo is a page: the definition of minimally qualified, the method and the panel, the data they saw, the trade accepted in Figure 2's terms, the within-band rule if there is a band, the review date.
The line also needs a review cadence, because everything beneath it moves. Applicant pools shift, roles change, and the norms that give a score its meaning age; the companion briefing on what test scores mean traces how reference populations drift, and a pass mark set on top of a norm inherits every year of that norm's age. Reviewing the mark on a schedule, against pass-rate trends and near-mark outcomes, treats it like the re-normed instrument it depends on: maintained, dated, and owned.
Handled this way, the meeting from the opening still happens, and the mark at 70 may still say no to the candidate at 69; no line can be kind to everyone a point below it. What changes is the sentence after the question. Instead of silence, there is a file: who placed the line, on what definition of the work, with which trade accepted, and when it comes up for review. The candidate gets the same verdict either way. The organization gets to keep it.
Where 5Profiler stands
A defensible mark needs a keeper. 5Profiler applies consistent, documented scoring rules identically to every candidate: pass thresholds are set per role, stored with the role's configuration, and enforced by the platform, so the rule one candidate met is the rule the next one faces, and every change to a threshold is a dated configuration event with an author. Adaptive testing supplies the ability scores beneath those thresholds, and the platform keeps the mark and its record in one place, which is most of what a serious review asks to see. The line still has to be chosen by people. What the platform guarantees is that the choice, once made, is applied and remembered.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
- Angoff, W. H. (1971). Scales, norms, and equivalent scores. In R. L. Thorndike (Ed.), Educational measurement (2nd ed., pp. 508–600). Washington, DC: American Council on Education.
- Cascio, W. F., Outtz, J., Zedeck, S., & Goldstein, I. L. (1991). Statistical implications of six methods of test score use in personnel selection. Human Performance, 4(4), 233–264.
- Cizek, G. J., & Bunch, M. B. (2007). Standard setting: A guide to establishing and evaluating performance standards on tests. Thousand Oaks, CA: Sage.
- Equal Employment Opportunity Commission, Civil Service Commission, Department of Labor, & Department of Justice. (1978). Uniform guidelines on employee selection procedures. Federal Register, 43(166), 38290–38315.
- Kane, M. (1994). Validating the performance standards associated with passing scores. Review of Educational Research, 64(3), 425–461.
- Livingston, S. A., & Zieky, M. J. (1982). Passing scores: A manual for setting standards of performance on educational and occupational tests. Princeton, NJ: Educational Testing Service.
- Society for Industrial and Organizational Psychology. (2018). Principles for the validation and use of personnel selection procedures (5th ed.). Industrial and Organizational Psychology, 11(Suppl. 1), 1–97.
- Taylor, H. C., & Russell, J. T. (1939). The relationship of validity coefficients to the practical effectiveness of tests in selection. Journal of Applied Psychology, 23(5), 565–578.