Research Scores & decisions

Every score wears a band.

At reliability .90 a 95% confidence band spans about ±0.6 SD; at .70, about ±1.1 — and near a threshold the band becomes the probability the verdict is wrong. The psychometrics every report reader is owed.

The report says 6.4. One decimal place of exactness, printed in the same confident type as the candidate's name, on a ten-band scale. Run the scale's own published reliability through the standard formula and the exactness dissolves. At a reliability of .80, a 95% confidence interval around that score spans roughly 4.6 to 8.2 in sten units (computed); the true value could plausibly sit anywhere from 5 to 8. The decimal is a claim about score precision that the measurement error underneath it cannot support.

Nothing went wrong in producing that number. The instrument was administered properly, the scoring was exact, and the software did what it promised. The gap between 6.4 and "somewhere between 5 and 8" has a plainer source: no test measures without error, and the size of the error is governed by test reliability, the proportion of variation in scores that reflects real differences between people rather than noise. Reliability is published in every serious manual. The calculation that turns it into a band takes one line. What most reports lack is not the science but the habit of printing the band beside the score.

The band matters because decisions are made at its edges. Its width follows from reliability once the scale is fixed; near a threshold, it converts into the probability that a pass is truly a fail; and error of this kind also pervades the supervisor ratings that validation studies score themselves against. An executive who absorbs one formula and a few published values will not out-model any vendor, and does not need to. The useful skill is smaller: knowing that every reported score is an estimate, and asking how wide the interval around the estimate is before letting the decimal settle anything.

Reliability sets the width of the band

Classical test theory, the accounting framework behind conventional scoring, starts from one identity: every observed score is a true score plus an error term. The error term collects everything transient in a testing moment: which items happened to be sampled, how the morning went, where attention drifted, which coin-flip judgment calls broke which way. A second sitting redraws all of it, which is why a person retested rarely lands on the identical number.

The size of that error term has a name and a formula. Harvill's (1991) instructional module for the National Council on Measurement in Education, written to teach precisely this to non-specialists, gives it in one line: SEM = SD × √(1 − reliability), where SEM is the standard error of measurement, the typical distance between the score a person obtained and the error-free score the test was estimating. Roughly 68% of observed scores fall within one SEM of the true score (Harvill, 1991), and widening the interval to about two SEMs yields the familiar 95% confidence interval, the range within which the true value plausibly lies.

The formula turns published reliabilities into bands, and the bands widen in strict order. At a reliability of .90, the SEM is 0.32 standard deviations and the 95% band spans ±0.62 SD. At .80, the SEM is 0.45 SD and the band is ±0.88 SD. At .70, the SEM is 0.55 SD and the band is ±1.07 SD (all computed). In sten units, at two stens to the standard deviation, those SEMs are 0.63, 0.89, and 1.10 stens (computed). Figure 1 renders the three bands to scale. The .70 row deserves a slow look: its full interval spans more than four sten bands, over 40% of the ten-band scale (computed). What a sten band is, and how percentile ranks stretch and compress along the scale, is the territory of the companion briefing on what test scores actually mean.

The progression is not linear: moving from .90 down to .70 costs 20 points of reliability but widens the band by nearly three-quarters (computed). Nor is any band narrow. Even at .90, the kind of reliability few scales exceed, the interval spans more than a full sten. A report can respond to this in two ways: measure with less error, or disclose the error it has. How adaptive tests pursue the first, steering item selection to shrink the standard error where it matters, is the subject of the companion briefing on adaptive testing in hiring. This article is about the second, because disclosure is the part every organization controls regardless of what it administers.

Disclosure is also, formally, the rule. The Standards for educational and psychological testing (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014), the profession's consensus reference, direct that score reports include indicators of precision such as the SEM or confidence bands, and that reliability evidence be reported for each score a user is expected to interpret. A report that prints a bare 6.4 is omitting information the field's own standards say the reader is owed.

Reliability is a more modest promise than validity: a reliable scale agrees with itself, so a retest lands close to the first sitting, while validity concerns whether the number tracks anything that matters for the job, a question with its own evidence base surveyed in the series' flagship review of what predicts job performance. A scale can be reliable about nothing. The reverse failure is the one this article turns on: an unreliable scale cannot be valid about anything, because a number that will not repeat cannot follow a stable criterion. Reliability is the ceiling under which every other claim a score makes has to fit.

Published reliability is a range, and facet scales sit at the wide end

Real instruments, when they follow the disclosure rule, publish a spread rather than a single figure. The professional manual for the NEO PI-R, the reference inventory of the five-factor tradition, is the instructive case because it prints its numbers plainly. The five domain scales report internal consistencies, the degree to which a scale's items agree with one another, between .86 and .92. The 30 facet scales beneath them report internal consistencies between .56 and .81 (Costa & McCrae, 1992). One instrument, one manual, and a spread of reliabilities wide enough to change what its scores mean.

Those values map directly onto Figure 1. A domain scale lives near the .90 row: a band of roughly ±0.62 SD, tight enough that a high score and a low score genuinely exclude each other. A facet in the middle of its published range sits near the .70 row and carries a band of ±1.07 SD. The floor of the facet range lies below every row Figure 1 plots, which means bands wider still. Two facet scores that look decisively different on the page can therefore be separated by less than the width of either score's interval, a possibility a reader cannot evaluate unless the interval is shown.

This is not an argument against facet scores, and the distinction matters commercially because facet-level detail is precisely what modern assessment sells. Narrow scales run on fewer, more homogeneous items, so they trade internal consistency for specificity: they measure a sharper construct with a wider band. The case that hiring decisions need that sharpness, that a domain average conceals the configurations roles actually reward, is made in the companion briefing on facet-level personality measurement and is not weakened by anything here. The conclusion is about reporting, so honest facet reporting shows the band: finer resolution and wider uncertainty arrive together, and a report that displays the first while hiding the second is editing the physics.

The criterion is noisier than the predictor

Reliability questions are usually aimed at the test, yet the noisiest number in the whole hiring system sits on the other side of the validation equation. When researchers estimate whether an assessment predicts job performance, "job performance" is nearly always a supervisor's rating, and ratings are measurements too, with error of their own. Viswesvaran, Ones, and Schmidt (1996) pooled the accumulated evidence on the reliability of job performance ratings in the Journal of Applied Psychology, and their central estimate has become the field's reference point: the interrater reliability of a single supervisor's rating of overall job performance is about .52. Ask two supervisors to rate the same people independently and barely half of what either produces is signal the other can see.

The geometry in Figure 2 is uncomfortable to look at. The criterion marker sits to the left of everything the personality manual publishes, below even the floor of the facet range. Organizations that interrogate a test's precision before buying it almost never interrogate the precision of the ratings they will later use to judge whether the test worked, and the second number is the worse one.

Error on either side of a correlation has a mathematical consequence: attenuation. Schmidt and Hunter (1996), working through 26 recurring research scenarios in Psychological Methods, showed how measurement error systematically shrinks observed relationships below the relationship between the underlying constructs, and how ignoring that fact misleads the conclusions drawn from data. Corrections exist and matter. For an executive the practical reading is concrete. A validity coefficient computed against noisy ratings understates what the assessment actually predicts, so the modest-looking coefficients in validation research are, if anything, conservative about the underlying relationships. And any internal list of "top performers" built from single-rater judgments is partly a record of who was rated, along with who performed.

Near a cut score, measurement error becomes flip risk

Bands become consequential at thresholds. The moment a continuous score meets a cut, a shortlist line, a progression gate, a minimum for interview, the question stops being "what is the score" and becomes "which side of the line is the true score on." Under the normal error model, the model classical test theory supplies for exactly this purpose, the answer depends on one quantity: the observed score's distance from the cut, measured in SEM units. An observed score half an SEM above the line has roughly a 31% chance that the true score sits below it. At one full SEM the chance is roughly 16%; at 1.5 SEMs, roughly 7% (all computed). And a score exactly at the cut is even odds, a verdict decided by a coin the instrument tossed.

The curve in Figure 3 falls quickly, and that is the reassuring half of the story: verdicts far from the line are safe, and most candidates are far from any given line. The unreassuring half is where hiring attention concentrates. Contested decisions, appeals, and agonized debriefs cluster near thresholds by definition, which means they cluster exactly where the flip probabilities are largest. The risk also runs in both directions. For every candidate passed with a 31% chance of truly belonging below the line, some rejected candidate sits half an SEM under it with an equal chance of belonging above, and that candidate rarely gets a second look.

How much of this risk a given program carries is knowable in advance, and almost nobody computes it. The exposure depends on where the line sits relative to the pool: a cut placed where candidates are densest manufactures many near-line verdicts, while a cut out in a tail touches few, so identical instruments can yield very different flip exposure in different hands. A program that keeps records can simply count how many of last quarter's decisions fell within one SEM of its line; each of those carried a flip probability worth knowing about, whichever way the verdict went.

These probabilities do not make cut scores illegitimate. They make a cut's placement and the band around scores near it one conversation, not two: a line defensible for a wide-band scale may sit somewhere quite different from a line defensible at domain-level precision. How to place and justify the line itself, from criterion-referenced standards to workforce arithmetic, is the subject of the companion briefing on setting cut scores. The obligation this article adds is narrower: whoever sets the line should know the SEM of the scale it is drawn on, in the scale's own units, and should treat that number as part of the line's documentation.

Why extreme scores drift toward the middle on retest

Classical theory makes one more prediction executives encounter without recognizing it. When an organization selects the highest scorers from an assessment round and tests them again, their average will drop, and the cause is the selection itself rather than any change in the people. An extreme observed score is more often than not a score that error helped: among everyone who posted a 9, favorable error outnumbers unfavorable error, since more true 8s got lucky than true 10s got unlucky. On a second sitting the error redraws at random while the true scores stay put, so the group's average slides back toward the mean. The drift is regression toward the mean, another face of measurement error, and it follows from the score-plus-error accounting behind the SEM (Harvill, 1991; Schmidt & Hunter, 1996).

Two managerial misreadings follow from not knowing this. The first treats a retest decline as evidence of decline, or worse, of first-round inflation: a candidate who scored 9 at screening and 8 at verification has done nothing suspicious, and a policy that flags such drops as anomalies will mostly flag statistics working normally. The second misreading runs the other way. Any ranked list is enriched with favorable error at its top, so the champions identified by one noisy measurement, a single test, a single rater's judgment, will on average look a little less exceptional on every subsequent measurement. Programs built on "hire only the 9s" quietly promise a precision the top of a ranked list never has.

Regression explains a pattern in performance data too, because supervisor ratings drift for identical reasons. An employee of the quarter is selected partly on favorable error in a rating whose interrater reliability sits near .52, so the next quarter's rating tends to land closer to the pack even when nothing about the person's work has changed. Expecting that pullback keeps a manager from reading ordinary regression as decline, and from crediting whatever initiative happened to be running when the statistics reverted on schedule.

The band belongs on every score a decision touches

The implications are reading habits, and none requires a statistician on staff. Start with the indispensable request: reliability for each reported scale rather than one headline coefficient for the instrument. The Standards require per-score reliability evidence, and the profession's own application of them to hiring, the SIOP Principles, carries the same expectation into personnel selection specifically (Society for Industrial and Organizational Psychology, 2018), so the request is not exotic; a vendor who answers with a single proud alpha for a 30-scale report has answered a different question, because an instrument's precision lives at the level of the scales a decision actually touches. The follow-up request matters as much: the SEM expressed in the metric the report displays, stens or percentiles or T-scores, because "the SEM is 0.45 SD" protects nobody in a debrief unless someone converts it on the spot.

The band's first job in day-to-day use is arbitrating comparisons. Two candidates whose scores on a scale differ by less than that scale's band are, for decision purposes, tied, and ranking within a tie is choosing by noise. This single habit dissolves a familiar meeting ritual: the earnest litigation of a 6.4 against a 6.9 on a facet whose interval is wider than the gap. When several scores are combined into one recommendation, the errors travel into the composite along with the scores; how that combination should be done, and by whom, is the subject of the companion briefing on mechanical versus holistic judgment.

Near thresholds, the band should harden into process. A verdict that falls within roughly one SEM of a cut carries a flip probability large enough that treating it as final is a policy choice, not a statistical one. The proportionate response is a review zone: scores inside the band around the line get a second source of evidence, a structured interview, a work sample, a supervised retest, before the verdict settles, on the same due-process logic the series applies to integrity flags. This is cheaper than it sounds, because the zone is narrow and most scores fall outside it. What it secures is the ability to say, to a candidate or a court or a board, that no one's outcome was decided by the coin inside the instrument.

The last habit is a preference: reports that print the interval over reports that assume it away. The disclosure already exists in every manual worth the name; the only question is whether it survives the trip to the page a decision-maker actually reads. You can see what surviving looks like in a sample assessment report, and the measurement foundations behind the scales themselves are set out on our science page. A report that shows bands does something subtler than inform: it trains every reader who touches it to expect measurement error as part of a score's anatomy, the way an engineer expects a tolerance beside a dimension.

There is a cultural dividend at the end of this. Teams that read bands stop arguing about differences the instrument cannot certify and redirect that hour toward evidence that could actually separate the candidates, which is a better use of both the hour and the instrument. The decimal will keep arriving; software is built to print it. What changes is its authority, and the change is a promotion for the measurement, from oracle to witness: testimony with stated margins, weighed alongside other testimony, which is all a well-made instrument ever offered in the first place.

Where 5Profiler stands

Ask where the band is, and a 5Profiler report answers on the page: score bands print with the scores themselves, so the interval travels with the number into every shortlist discussion and debrief instead of staying behind in a technical manual. Scale-level precision is treated as reportable information, at the facet grain where bands are widest and over-reading is most tempting. Upstream, adaptive testing manages the width of what will eventually be printed, selecting each next item to narrow the interval around a candidate's estimate. The decimal still arrives on a 5Profiler report; the band arrives beside it.

Read the science behind the platform · See it on your roles

References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
  2. Costa, P. T., Jr., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO PI-R) and NEO Five-Factor Inventory (NEO-FFI) professional manual. Odessa, FL: Psychological Assessment Resources.
  3. Harvill, L. M. (1991). Standard error of measurement. Educational Measurement: Issues and Practice, 10(2), 33–41.
  4. Schmidt, F. L., & Hunter, J. E. (1996). Measurement error in psychological research: Lessons from 26 research scenarios. Psychological Methods, 1(2), 199–223.
  5. Society for Industrial and Organizational Psychology. (2018). Principles for the validation and use of personnel selection procedures (5th ed.). Industrial and Organizational Psychology, 11(Suppl. 1), 1–97.
  6. Viswesvaran, C., Ones, D. S., & Schmidt, F. L. (1996). Comparative analysis of the reliability of job performance ratings. Journal of Applied Psychology, 81(5), 557–574.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.