Research Scores & decisions

A score is a comparison.

The same performance can read as ordinary or exceptional depending on the norm group beneath it — and norms age at a measured 2.31 points a decade. How to read any assessment number: compared with whom, measured when.

One candidate, one sitting, one performance, and a report that renders it three ways: the 79th percentile on the summary page, a sten of 7 on the profile grid, a T-score of 58 in the technical appendix. The numbers look like separate findings. They are a single finding in different units, and it is not a finding about the candidate alone: each one says that this performance sits roughly eight-tenths of a standard deviation above the average of a particular reference group. What test scores mean is settled by that group, the norm group, and the most consequential line on any report is the one that names it. Often it is in small print. Sometimes it is missing.

Measurement people call a score of this kind norm-referenced, after Glaser's (1963) distinction between scores that locate a person relative to other people and scores that certify performance against a fixed standard. Nearly everything on a talent-assessment report is the first kind. A percentile rank, a sten, a T-score: each exists only once a comparison group has been chosen, because each is a statement about position inside that group's distribution. A test score is not a property of a person; it is a property of a comparison. Change the comparison and the number moves, while the candidate stays put.

That dependence is not a technicality. Norm groups differ, so one raw performance can read as ordinary against one table and exceptional against another. Percentile units stretch and compress along the scale, so gaps that look equal on the page are not equal amounts of anything. And norms age: against a fixed table, scores drift upward at a meta-analytic 2.31 standard-score points per decade (Trahan, Stuebing, Fletcher, & Hiscock, 2014), which means an old norm inflates every cohort measured after it. Each of these problems yields to the question this briefing keeps asking, compared with whom and measured when, and a defensible program is one that can answer it for every number it prints.

Percentile units compress where most candidates sit

Vendors report position on a family of interchangeable scales, and the family is worth a paragraph of anatomy. A z-score counts standard deviations from the norm group's mean. A T-score rescales z to a mean of 50 and a standard deviation of 10, clearing away negative signs and decimals. Sten scores, short for standard ten, compress the range into 10 bands with a mean of 5.5 and a standard deviation of 2, a format fixed in the handbook of the 16PF (Cattell, Eber, & Tatsuoka, 1970) and still standard across personality reporting. A percentile rank states the share of the norm group scoring below the candidate. Figure 1 aligns them all beneath one normal curve, with the report from the opening paragraph marked as a single dashed line crossing every scale. Every conversion printed in this briefing, including that one, is computed under the normal model rather than taken from any dataset.

If they all encode one position, why keep them all? Convention and audience. The z-scale suits people who think in standard deviations. The T-scale spares everyone else the minus signs. Sten scores resist over-reading by design, which the closing sections will return to. Percentiles are vivid for lay audiences and survive every committee meeting for that reason. The plurality is harmless as long as a reader treats the scales as notations for one comparison. It stops being harmless when the notations get mistaken for corroboration, a report reader nodding at the T-score and the percentile as if two independent facts agreed.

The percentile is the ruler lay readers reach for, and it is also the one that distorts distance. A normal distribution piles people near its middle, so a move of fixed size in trait units crosses a crowd at the center and nearly nobody in the tail. In sten terms: the central band, sten 6, spans roughly 19 percentile points, while the top band, sten 10, spans about 2 (computed). A single step of identical width in trait units reads as a leap in the middle of the scale and as a rounding error at the end of it, which is what the brackets in Figure 1 record.

Textbook psychometrics states the property without hedging: percentile ranks are ordinal, and equal percentile differences do not correspond to equal differences in raw score or in the underlying trait (Crocker & Algina, 1986). Ranks preserve order and discard distance. The practical casualties are the small habits of report arithmetic: subtracting one candidate's percentile from another's, averaging percentiles across scales, setting thresholds at round percentile values. Each performs addition on a scale built for ranking.

The sturdier habit runs in the other direction: hold a candidate's position in standard-deviation units, then let the percentile illustrate it for intuition. The sten and the T keep their spacing because they stay in standard-deviation units; the percentile is the translation into crowd language, vivid and nonlinear. Read from the standard scale outward, and treat any sentence built on percentile subtraction as a sentence about ranks.

The norm group decides what test scores mean

Scales are the visible machinery. The load-bearing choice sits underneath them: which distribution the candidate is being placed in. A vendor's norm table is a sample of prior test takers, perhaps a general adult sample, perhaps graduate applicants, perhaps experienced managers in one sector. Figure 2 stages the comparison that report readers rarely get to see. One raw performance, held fixed, lands at the 93rd percentile against a general-adult norm, the 76th against a graduate-applicant norm, and the 62nd against an experienced-manager norm (illustrative). Nothing about the candidate varies between the panels. The comparison varies, and the number follows it.

All of those percentiles are true at once, which is exactly the difficulty. A recruiter who tells a hiring committee 93rd percentile has made a claim whose content depends entirely on a table the committee will never see. Attached to a defined group, the number informs. Detached from the group, it drifts toward a compliment. The percentile's fluency is what makes the detachment easy: it sounds like a property of the person, the way height is, and the resemblance is the whole danger.

The profession's rulebook anticipates this. The Standards for Educational and Psychological Testing require norms to describe defined populations, hold norming samples to representativeness, recency, and relevance, and expect score reports to support accurate interpretation (American Educational Research Association et al., 2014). The International Test Commission's guidelines assign the duty where it lands in practice: the test user must select norms appropriate to the population being assessed and to the purpose of the assessment (International Test Commission, 2001). Choosing the norm group, in other words, is part of test use, on the same footing as choosing the test.

The choice is substantive because different norm groups answer different questions. An applicant norm answers: how does this candidate compare with the people competing for roles like this one? An incumbent norm answers: how does this candidate compare with the people already doing the work? Both are legitimate questions, they are simply different ones, and a selection screen usually calls for the first, because the decision being made is a choice among applicants. Benchmarking candidates against incumbents also imports an incentive gap: applicants score higher than incumbents on desirability-loaded scales, by about half a standard deviation on conscientiousness, a pattern this series unpacks in its briefing on faking, so part of a percentile computed against an incumbent table reflects the incentive difference rather than the trait.

Norm-group identity is also where sales language is at its vaguest. A phrase like top quartile on our global professional benchmark performs rigor while naming no population, no sample size, and no norming date. The Standards' vocabulary is the antidote a nonspecialist can carry into any vendor meeting: defined population, representativeness, recency, relevance. A norm that cannot be described in those terms cannot be interpreted in any terms, and a percentile computed on it is a number without a referent.

Sample size and recruitment matter for the reason they matter in polling: small or self-selected norm samples make unstable reference points, and every percentile read against them inherits the instability. Representativeness is the harder obligation, because a norm built from whoever happened to take the test describes that accident of recruitment, and a large accident is still an accident. Organizations with enough assessment volume sometimes commission local applicant norms for exactly this reason: the reference group becomes the organization's own applicant flow, and compared-with-whom acquires its sharpest possible answer, at the price of upkeep that most programs underestimate.

A single report can also carry several norm groups at once. Multi-instrument batteries commonly norm each instrument on whatever sample its development history provided: the personality scales against one reference pool, the ability scales against another, collected in different years from different populations. Nothing requires those crowds to match, and often they do not. The percentiles then sit in adjacent columns as if they shared a reference group, and a reader comparing across columns is quietly comparing across crowds. Where the groups differ, that fact belongs on the report, and on the reader's list of things to check when a cross-scale contrast looks striking.

Test norms age at a measured rate

Even a well-built norm has a shelf life, because populations move under fixed instruments. Flynn (1984) documented the pattern that now carries his name: massive gains on American IQ tests scored against fixed norms, roughly 13.8 points between 1932 and 1978. The instruments had stood still while the people taking them changed, and each re-norming revealed how far the old table had fallen behind the population it claimed to describe.

Trahan, Stuebing, Fletcher, and Hiscock (2014) gave the drift its modern estimate in a meta-analysis published in Psychological Bulletin, spanning 285 comparisons and roughly 14,031 test takers: scores rise against fixed norms at a mean of 2.31 standard-score points per decade, and at around 2.93 per decade on modern Wechsler and Stanford–Binet batteries. Why populations drift this way remains unresolved, and this briefing takes no position on it. The consequence does not depend on the cause: a fixed table gradually comes to sit below the population it describes, and every score read against it inflates.

Figure 3 turns the rate into a maintenance chart. At 2.31 points per decade, a norm table left unrefreshed for 15 years reads a cohort roughly 3.5 standard-score points high, about a third of a standard deviation (computed from the meta-analytic rate); on the steeper modern-battery slope the drift accumulates faster still. The inflation is invisible from inside any single report, because every candidate is lifted together and rank order within the cohort is undisturbed. It surfaces at the joints: when this year's cohort is compared with one tested a decade earlier, when a pass mark set against fresh norms is reused against stale ones, or when a report presents a percentile as if the reference crowd were contemporary.

Norm currency is therefore a property a buyer can audit with a single question: when were these norms collected? The date belongs beside the norm group's name in every technical manual, and a vendor unable to produce it is reporting positions inside a crowd of unknown age. The audit requires no statistical training, only the habit of treating test norms like other reference data, salary surveys, market benchmarks, actuarial tables, none of which anyone would quote a decade stale without comment.

Re-norming itself has a visible signature, and knowing it prevents a misreading. When a vendor refreshes norms after a long gap, raw performances identical to last quarter's suddenly land at lower percentiles, because the reference crowd has caught up. Recruiters experience this as the test getting harder. It is the yardstick getting truer, and a program that understands Figure 3 will read the discontinuity as maintenance rather than malfunction. The communication plan writes itself once the mechanism is understood: announce the refresh, restate the norm group, and re-anchor any standing thresholds against the new table.

The report is coarse on purpose

Seen against all of the above, the coarseness of the classic personality report is a design decision with reasons behind it. Sten scores can say only that a candidate sits in band 7 of 10; they decline to distinguish a 7.2 from a 7.4, and the refusal is deliberate, because distinctions finer than the instrument can support would be noise typeset as information. This briefing has treated every printed score as exact, and no real score is. The width of the uncertainty band around a number, and the way that band can flip a decision near a threshold, is the subject of the companion briefing on measurement error; alongside compared-with-whom and measured-when, how-precisely completes the set of questions a reader owes any printed score. This briefing stays with the comparison, because even an error-free score would mean nothing without one.

Coarseness also shapes the conversation a report can support. Sten bands invite a committee to discuss whether a candidate stands broadly high, middling, or low on a trait, which is a claim the instrument can stand behind. Two-decimal percentiles invite rank-ordering of near-identical candidates, which it cannot, and the false precision migrates straight into meeting minutes. A band is a fence against over-interpretation, placed by the people who know the instrument's limits best.

Reports also mix Glaser's two reference systems on a single page, and much confusion about what test scores mean is a confusion between them. The sten, the T, the percentile: norm-referenced, all of them. Sentences about a role's requirements are criterion-referenced; they compare the candidate with a demand rather than with a crowd. The systems fail differently. A norm-referenced line inherits every weakness of its norm group, while a criterion-referenced line inherits every weakness of the role analysis behind it, so a careful reader asks compared-with-whom of the first kind and asks who defined the demand of the second. How role demands come to be quantified at the grain of individual facets is the territory of the briefing on facet-level measurement.

The last reading habit is refusing to let relative standing impersonate behavior. A percentile rank says what share of a reference group scored lower. It does not say how often the candidate will speak up in a meeting, finish work ahead of schedule, or stay composed with an angry customer; claims about behavior rest on item content and validity evidence, and no norm group converts a rank into a frequency. A report reads best when its percentile column is treated as placement, its narrative text as hypothesis to be probed at interview, and its role-referenced sections as the point where placement meets a specific job. A sample report shows the anatomy in a single artifact, norm-referenced and criterion-referenced elements side by side.

A defensible program can name its comparison

The practical program is short. Every scale a vendor reports should arrive with its norm group's identity, its sample size, and its norming date, in the technical documentation and ideally on the report itself. The request is the Standards' own vocabulary put to work (American Educational Research Association et al., 2014), and it fits in one email. A vendor with well-kept norms answers it quickly; how scoring is assembled end to end is the kind of documentation a serious provider makes public, in the shape of a platform overview a buyer can actually read.

Matching the norm to the decision comes next. Selection screens compare applicants, so applicant norms usually fit the question being asked. Development programs work with incumbents, where an incumbent reference makes sense. Graduate schemes read most cleanly against graduate applicant pools instead of the general adult population. The reasoning mirrors what this series says about validity evidence: borrowed numbers owe you a check on whether they transport to your setting, the subject of the local validation playbook. A percentile estimated on someone else's sample is borrowed evidence too, and the norm-group question is the transport check it has to pass.

Re-norming belongs on the maintenance calendar with everything else that decays. The measured drift gives the schedule a floor: at 2.31 points per decade, a decade of neglect is a visible bias, and the 15-year marker in Figure 3 is the case for never letting the question go unasked that long. Contracts can carry the obligation, with the vendor stating the norming date, notifying on re-norms, and flagging when a scale's norms pass an agreed age, the way other reference data carries an as-of date.

Train report readers last, because the training is what makes the rest stick, and because what test scores mean reduces in practice to the questions a reader remembers to ask. Compared with whom: name the norm group before repeating any percentile to a committee. Measured when: check the norming date, and discount accordingly. How precisely: put the error band around the number before it decides anything close, as the measurement-error briefing sets out. A debrief that opens this way sounds pedantic exactly once. After that it sounds like the house style of a program that knows what its numbers are.

The habit pays off longest after the decision, because scores outlive their moments. A percentile earned in this spring's screen resurfaces in a talent review years from now, quoted from memory, its norm group and norming date long detached, its cohort aged against whatever table now sits in the system. Files travel; context stays behind unless someone writes it down. A program that records compared-with-whom and measured-when beside every number it stores is protecting more than the original decision. It is protecting every later reader of the file, including the ones who will never think to ask.

Where 5Profiler stands

Compared with whom, measured when: a 5Profiler report is built so the answers sit on the page. Where a candidate is read against a role's demands, the report names that as role-referenced scoring; where a number is norm-referenced, the norm group and the recency of its norms are stated beside the score, so no percentile travels without the reference that gives it meaning. The transparency holds at 30-facet personality resolution, facet by facet, and it holds in the archive: a colleague who opens a report years after the decision inherits the comparison along with the number, and the file keeps answering the questions its first reader asked.

Read the science behind the platform · See it on your roles

References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
  2. Cattell, R. B., Eber, H. W., & Tatsuoka, M. M. (1970). Handbook for the Sixteen Personality Factor Questionnaire (16 PF). Champaign, IL: Institute for Personality and Ability Testing.
  3. Crocker, L., & Algina, J. (1986). Introduction to classical and modern test theory. New York: Holt, Rinehart and Winston.
  4. Flynn, J. R. (1984). The mean IQ of Americans: Massive gains 1932 to 1978. Psychological Bulletin, 95(1), 29–51.
  5. Glaser, R. (1963). Instructional technology and the measurement of learning outcomes: Some questions. American Psychologist, 18(8), 519–521.
  6. International Test Commission. (2001). International guidelines for test use. International Journal of Testing, 1(2), 93–114.
  7. Trahan, L. H., Stuebing, K. K., Fletcher, J. M., & Hiscock, M. (2014). The Flynn effect: A meta-analysis. Psychological Bulletin, 140(5), 1332–1360.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.