Research The AI era

The grader that changed overnight.

On benchmark essays ChatGPT agrees with human raters inside the human range — and the same product’s accuracy on one bounded task fell from 97.6% to 2.4% in three months. The governance that makes machine grading usable.

The pitch for an LLM grader is stamina. It reads ten thousand open-ended answers before lunch, gives the ten-thousandth the attention it gave the first, and applies the rubric without a Friday-afternoon mood or a favorite kind of candidate. LLM scoring, in other words, promises the thing human rating pools have never delivered: consistency at scale. The catch arrives on a version number. A grader built on a commercial model can change overnight, on the provider's schedule, and the system that scored last quarter's applicants and the system scoring this quarter's may share nothing but a name.

Machines grading open text is not a ChatGPT-era experiment. The founding proposal is from 1966, the operational track record runs back decades, and the measurement field long ago wrote down how any machine grader should be judged: against trained human raters, task by task, group by group, and continuously after deployment. That framework converts a culture-war question, whether software should grade writing, into an empirical one: where, exactly, does the agreement hold?

Held to that standard, the new graders split cleanly. On benchmark essays, ChatGPT's holistic scores agree with trained humans at levels inside the human range; its trait-level scores fall well short of it; short written answers are shakier still. And the new graders can fail in a way the operational era never saw, because the model behind them can be revised mid-cycle with no action from the assessment owner. The conclusion the evidence supports is not "never." It is "under governance": frozen versions, a human-scored calibration set, sampled double-scoring in production, and an appeal that ends with a person.

Machine grading is decades older than the chatbot

The founding document announces itself. Ellis Page's 1966 article in Phi Delta Kappan was titled "The imminence of grading essays by computer," and the title carried the thesis (Page, 1966). The prediction ran ahead of the technology by some decades, but it landed. Automated essay scoring, software that assigns scores to written responses, spent the years after Page maturing from proposal to production.

Attali and Burstein (2006) mark where it landed: e-rater V.2, an automated essay-scoring engine in operational use in testing programs, with machine–human agreement comparable, in its deployments, to the agreement between two trained human raters. For the current debate, the engineering matters less than the posture around it. The engine's owners published how it was evaluated, showed its agreement against human judgment, and treated a consequential score as a claim requiring evidence, because the programs it scored were real and contested.

A demonstration and a deployment make different promises. A demonstration shows that an engine can score a stack of essays once. An operational program commits to scoring every response, all cycle, with the results feeding real decisions, and it accepts the obligations that follow: published evaluations, human raters kept in the loop, and procedures for the responses the engine reads badly. Automated essay scoring reached operational status program by program, evaluation by evaluation, and the grammar of those evaluations is the subject of the next section.

That history reframes what LLM scoring actually has to prove. Producing a number was settled two decades ago; software has scored operational essays for that long. The live question, then and now, is whether the number agrees with disciplined human judgment, and whether the agreement survives new prompts, new populations, and time. One scope note before the evidence: this briefing stays on text, meaning essays, short answers, and written exercises. Spoken answers, where a system transcribes and scores what a candidate says, carry extra measurement layers and have their own briefing on AI-scored interviews.

A machine grader is qualified the way a human rater is

The rulebook was codified by Williamson, Xi, and Breyer (2012), measurement scientists writing in Educational Measurement: Issues and Practice, as a general framework for the evaluation and use of automated scoring. It treats the machine as a rater and asks of it what rater-training programs ask of people: whether the engine agrees with trained human raters; whether that agreement matches the agreement between two trained humans scoring the same responses; whether it holds across the subgroups the test reaches; whether it degrades over time or on new prompts; and whether there are defined lanes in which a human scores instead, for the cases the machine should not touch.

Built into the framework is an assumption that automation is partial. Some responses are off-topic, adversarial, or simply unlike anything the engine was evaluated on, and the framework expects a scoring program to define, in advance, which responses get routed to a human reader instead of scored by the machine. A program that cannot say what its machine should not score has not finished describing its machine.

Agreement in this literature has a standard currency: quadratic weighted kappa, or QWK, an agreement statistic for ordered ratings that credits near-misses and runs from 0 at chance to 1 at perfect agreement. The consequential move in the framework is the baseline. Open responses have no answer key, so the machine is judged against trained human raters, who disagree among themselves; the target worth hitting is the band that human pairs occupy. What rater disagreement of any kind does to a decision downstream is band machinery, covered in the companion briefing on measurement error.

The testing Standards put an institutional floor under the framework. Scores must remain comparable over time and across conditions of administration, and the scoring procedure counts as part of the instrument, owing evidence of its own (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014). Neither document is anti-automation; e-rater passed this style of scrutiny inside operational programs. The machinery is a refusal to let "the software scored it" end the conversation, and it applies to every novel scorer, including the telemetry-derived scores examined in this batch's review of gamified assessments.

This briefing quotes none of the framework's operating numbers on purpose. What transfers across programs is the logic, comparisons to run and monitoring to sustain, with thresholds set by each program's stakes and rating design. That blocks a shortcut both enthusiasts and skeptics reach for: there is no number to quote from someone else's program, only an engine's measured agreement on your tasks, against your raters.

Where LLM scoring clears the bar, and where it misses

The cleanest current test of the new graders comes from Shermis (2025) in Assessing Writing, which ran ChatGPT through the framework's opening comparison on the ASAP corpus, publicly released essay sets scored by trained human raters and long used to benchmark automated essay scoring. Scored holistically, one overall judgment per essay, ChatGPT agreed with the human ratings at QWK .67 to .84 across the essay sets. Human–human pairs across the program ranged from .61 to .97. The machine's band sits inside the human band: above the weakest trained pairs, below the strongest.

The picture inverts when the grading gets specific. Asked for analytic scores, ratings of individual rubric dimensions instead of one overall judgment, ChatGPT's agreement fell to QWK .18 to .63, below the human raters; the floor of that range sits closer to chance than to the weakest human pairing in the program. Short-form constructed responses, the sentence-to-paragraph answers most hiring exercises are built from, resisted summary altogether: performance was inconsistent across items, too unstable to state as a range.

Set on one axis (Figure 1), the three bands give the honest summary its shape: holistic essay scoring is at the bar; trait-level and short-form scoring are not. The split runs opposite to hiring's needs. Employers rarely need a ranked pile of five-paragraph essays; they need short answers scored against specific rubric dimensions, written case responses read for defined competencies, and open-ended scenario judgments, the scoring-key territory reviewed in the briefing on situational judgment tests. Scoring open-ended responses at hiring's grain is exactly where the current agreement evidence is weakest.

One reading of the split is mechanical. A holistic score can succeed by aggregation, many weak signals about vocabulary, development, and structure pooling into one stable overall judgment, while an analytic score forbids the pooling and asks the model to isolate a single dimension, and a short answer offers little text to pool in the first place. Whatever the mechanism, the operational rule survives it: agreement is a property of the task shape, and a validation run on essays licenses essay scoring and nothing else.

The human band is itself informative. Trained, monitored raters on a mature program agreed anywhere from .61 to .97, task to task; human scoring at scale is a managed process with real variance in it. Organizations that recoil from machine scores as unnatural are usually comparing them to a remembered human grader who never existed. The right comparison is the machine against the raters you actually run, on the material you actually score.

Read as a deployment guide, the evidence licenses a narrow configuration: holistic scores on extended writing, with everything else human-scored or human-anchored. That is a real capability. But it is a fraction of the open-response scoring a hiring funnel wants to buy, and a vendor demonstration that shows essay agreement while the product scores short answers is offering evidence from the wrong row of the figure.

A moving model breaks score comparability

e-rater changed only when its owners changed it. A grader rented from a model provider inherits different physics: the engine behind the interface can be revised on the provider's schedule, and the assessment program does not run the release calendar. Chen, Zaharia, and Zou (2024) measured what that means, in Harvard Data Science Review, by sampling two widely used commercial models, GPT-3.5 and GPT-4, in March 2023 and again in June 2023, on the same task families.

The task worth plotting is the simplest one: decide whether a number is prime, a question with one right answer and no room for taste. In March, GPT-4 got it right 97.6% of the time. In June, the figure was 2.4%. Across the identical window, GPT-3.5 traveled the other way, from 7.4% to 86.8%. Figure 2 plots the four points, and the crossing lines carry the finding: neither June model behaved like its March self. Nor was the movement confined to this one task; the authors documented substantial behavior change across multiple task families.

For measurement, the direction of the drift is beside the point. GPT-3.5's improvement breaches score comparability exactly as thoroughly as GPT-4's collapse, because the Standards' demand is that a score mean one thing across time and conditions of administration (American Educational Research Association et al., 2014). An unpinned grader fails that demand in both directions at once: two applicants for one role, scored three months apart, have been read by different raters who share a product name. The series makes a related argument about aging norms, that a score's meaning carries a timestamp, in the briefing on what test scores mean; with an unpinned LLM grader, the timestamp attaches to the grader itself.

The study's particulars will age; the models sampled in 2023 have successors already. The measurement lesson does not age with them, because it never depended on which way any given model moved. A grader whose behavior can change outside the program's control has a property no instrument is permitted to have, and absent governance the property is standing, built into how commercial model services ship. Assessment owners cannot assume the June grader is the March grader. They can verify it, or pin the version so the question never arises.

This briefing is about models grading candidates; the mirror image, candidates using models to write the answers, is a different problem with different economics, and it belongs to the companion briefing on AI-resistant assessment design. The drift result sits squarely on the grader side: between March and June, the candidates' task never moved. The scoring pipeline did.

Governance that makes a machine rater usable

The controls that answer this evidence come straight from the operational era; each restates the Williamson framework or the Standards for a grader that can move. Figure 3 connects them into a loop. Two of them carry most of the load: the frozen version and the calibration set.

Version pinning means a hiring cohort is scored, start to finish, by one frozen model configuration. The commitment is comparability's minimum: everyone competing for the role meets the grader in the state it was qualified in, and if the configuration must change, the cohort boundary is where it changes. The calibration set is the discipline's memory: a stock of real responses scored by trained human raters, held constant, and rerun whenever anything in the scoring path changes. A model update, a prompt edit, or an engine swap triggers the rerun, and agreement against the human anchor is recomputed before any live candidate is scored. The prompt belongs in that sentence deliberately: in an LLM scoring pipeline, the prompt is part of the scoring procedure, and the Standards' logic makes it part of the instrument.

Production adds the standing checks. A sampled share of live responses is double-scored blind by trained human raters, and machine–human agreement is tracked as an operating metric, by task type and over time; this is the framework's degradation monitoring, run forever rather than until comfort sets in. The framework also prescribes agreement checks by subgroup, and the prescription stands even though this briefing quotes no subgroup findings: a grader that has slipped for one group of candidates should be caught by the program's own monitoring, before anyone outside it notices (Williamson, Xi, & Breyer, 2012).

Two controls face the candidate. A consequential machine score must be explainable at the level the decision requires, a duty with its own briefing in this batch on algorithmic transparency. And a contested score needs an appeal lane that ends at a qualified human re-score, the due-process logic this series develops for integrity flags, applied to grading. Both depend on one artifact: the documented rubric, written before scoring begins. The rubric makes the construct, not the grader's taste, the thing being measured; it gives the human on appeal something to score against; and it turns a machine–human disagreement into evidence about the grader instead of a war of impressions.

The loop is format-general; neighboring formats add their own evidence. Where the artifact is code, execution changes the problem, since a submission can be scored against tests it passes or fails, and that literature is reviewed in the briefing on coding assessments in hiring. Text has no executable oracle. For essays, short answers, and written exercises, the rubric and the human anchor are the only ground truth available, and that is what makes the calibration set the control the others lean on.

Qualifying the machine rater never ends

For an organization weighing AI grading, the evidence assembles into a posture: an LLM grader is a rater you must qualify, not a feature you enable. Qualification is concrete. It means an agreement file, machine–human QWK by task type beside human–human agreement on comparable material, and on current evidence that file will read well for holistic essay judgment and poorly for trait-level and short-form scoring. A supplier who cannot produce the file is asking to seat an unqualified rater in front of your candidates. The file has a human side too: it names the raters the machine was compared against, the material they scored, and the date of the comparison, because agreement measured against last year's raters on last year's prompts is a dated fact, and Figure 2 is the argument that such facts age.

Contracts have to carry the service's physics. Version pinning, advance notice of model and prompt changes, and the right to rerun a calibration set before accepting a new configuration are ordinary clauses once the grader is understood as an instrument. The scoring procedure is part of the instrument on the Standards' own terms (American Educational Research Association et al., 2014), and a provider who can alter it mid-cohort has quietly taken over a piece of your validity argument.

Inside the organization, the assignment is to keep the human anchor funded. Trained raters do not disappear from a program that adopts machine scoring; they move upstream and sideways, into building calibration sets, double-scoring samples, and hearing appeals. The cadence never ends, because agreement is an operating metric of the program, and the drift evidence is the standing reason it cannot be treated as a launch-time metric alone. Budget for it the way rating pools were always budgeted for, as part of what scoring costs: machine speed changes the volume, and it does not change the obligation.

The posture also settles where open-response scoring belongs in a hiring design: alongside structured instruments, under documented standards, sized to what the current agreement evidence supports. A written exercise scored against a published rubric holds a defensible place next to adaptive cognitive measures and structured personality assessment on a modern assessment platform, and what a rubric-anchored score looks like when it reaches a decision-maker is visible in a sample candidate report. What the evidence does not yet support is delegation without an anchor: short answers and fine-grained rubric traits scored by an unmonitored model, at scale, with nobody rechecking the agreement.

Sixty years after Page called machine grading imminent, the machine part is a commodity. The scarce part is what the operational era built around its engines and what the drift result proves LLM scoring cannot skip: a standing comparison against disciplined human judgment. Ten thousand answers before lunch is a real offer, and on whole essays the agreement numbers say it can be taken. What the offer does not include is permanence. The grader can be tireless; the qualifying of it can never be finished.

Where 5Profiler stands

Open-response exercises enter a 5Profiler assessment with the scoring standard already on paper: human-anchored rubrics that record, level by level, what a scored answer looks like before any candidate writes a word. That documentation is the platform's answer to the question this briefing keeps asking of machine graders, because a written standard is what an agreement check needs to agree about, what a human re-scorer returns to, and what keeps the construct in charge of the score. Role-referenced scoring decides which open-response exercises a role warrants alongside its structured instruments. And whatever grades at scale, the anchor of record stays a human-scored calibration set: the standing exam any grader, human or machine, has to keep passing.

Read the science behind the platform · See it on your roles

References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
  2. Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater V.2. Journal of Technology, Learning, and Assessment, 4(3).
  3. Chen, L., Zaharia, M., & Zou, J. (2024). How is ChatGPT's behavior changing over time? Harvard Data Science Review, 6(2).
  4. Page, E. B. (1966). The imminence of grading essays by computer. Phi Delta Kappan, 47(5), 238–243.
  5. Shermis, M. D. (2025). Using ChatGPT to score essays and short-form constructed responses. Assessing Writing.
  6. Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 31(1), 2–13.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.