Research Applied playbooks

Sixty briefings. One operating rhythm.

The field’s own standard takes two documents to describe one assessment: the provider’s half and the client’s. This capstone draws the client half — the operating rhythm that keeps evidence-based hiring true after the purchase.

It took the field two documents to describe one assessment. ISO 10667, the international standard for assessment service delivery, is written in two parts: one sets requirements for the service provider, the party that supplies instruments, runs sittings, and returns scores; and one for the client, the organization that chooses assessment procedures, implements them, evaluates them, and answers for the data they generate (International Organization for Standardization, 2020a, 2020b). Evidence-based hiring is decided in the second part. The provider's half of the work can be delivered; the client's half can only be run, and it is the half organizations forget is theirs.

The forgetting is rarely a decision. An organization adopts validated instruments, structures its interviews, turns on proctoring, and reasonably concludes that it now practices what the research recommends. What it has installed is the provider's side of the ledger. The client's side, the choosing, the evaluating, the answering-for, never arrives in a box: it has no launch date, no go-live, and, in most organizations, no name attached. Nobody declined the work. Nothing on the org chart ever made it exist.

That gap is what this final briefing is for. The capstone of the first series set out the commitments a standards-based assessment program makes; this one is about the rhythm that keeps them kept. An operating model, in the sense this article intends, is everything about hiring that remains your work no matter how good the tooling gets. What makes it easy to forget is that nothing in the organization is named after it. There is no seat to be appointed to and no box on the chart to inherit, so the work goes on being defended in principle by people with no date on which to do it.

Evidence-based hiring decays by default

No meta-analysis has ever been overturned in a boardroom. Where evidence-based hiring fails inside organizations, the evidence is rarely what failed; the upkeep is. A cut line is placed at launch, with care, and never revisited while the applicant pool it was drawn against moves out from under it. Interview structure erodes one accommodation at a time, a favorite question here, an unscored improvisation there, until the panel is again running the conversation the briefing on structured interviews begins by burying. Bands drop off score reports in the name of tidiness, and decisions resume treating a point estimate as a fact, the habit the briefing on measurement error for executives exists to break. Norms and role profiles age in place. And quality of hire, the outcome the whole apparatus was adopted to improve, goes unrecorded, so no part of it registers anywhere as a result.

These failures share an anatomy. Each artifact was built by someone whose job was building it and is maintained by no one, because maintenance appears on nobody's objectives. Adoption is an event, so it attracts an owner, a budget, and a date. Upkeep is a condition, so it attracts goodwill. Decay is the default state of a hiring program for the plainest organizational reason there is: entropy is unassigned.

The field's documents assume none of this will be left to chance. The SIOP Principles treat the evaluation of selection procedures as an ongoing professional expectation, part of what using them responsibly means (Society for Industrial and Organizational Psychology, 2018). And the client part of the service-delivery standard is blunt about where the wider duties live: choosing assessment procedures, implementing them, evaluating them, and handling participant data appear there as requirements on the organization being served (International Organization for Standardization, 2020a). The documents are not describing a vendor's product line. They are describing your calendar.

A decayed program, meanwhile, still runs, which is what makes decay safe from attention. Assessments go out, scores come back, offers get made, and every quarter resembles the last. The program keeps producing decisions long after it has stopped producing evidence about them, and no alarm separates the two conditions, because the difference lives in artifacts nobody is scheduled to open. Nothing about decay announces itself. Prevention has to live on a schedule someone owns, and building that schedule is the rest of this article.

The four quadrants a hiring program has to keep turning

The hiring operating model this series closes on is deliberately small, and Figure 1 holds the whole of it on one standing cadence. Define what the funnel selects for. Run the process as designed. Read the evidence with its controls on. Learn what the hires became, and let the learning revise the definitions. Each quadrant produces artifacts the next one consumes, which is what makes the model a cycle and its absence quiet: a program can skip a quadrant for years and feel nothing, because missing artifacts are only ever missed downstream.

Definition, and the reasons attached to it

Definition is the quadrant organizations most often believe is finished, because its artifacts read like paperwork: a role profile stating what the role demands, a criterion definition stating what good performance will mean before anyone is hired against it, a cut policy stating where the lines sit. The working test is whether each artifact carries its rationale. A cut line with a written justification can be revisited when the pool shifts, the maintenance case the briefing on setting cut scores makes in full. A norm choice with a date on it can be checked for currency, the check the briefing on what test scores mean shows readers how to run. A criterion definition with a named owner can be argued with, which is what it is for. Definitions without rationales are settings, and settings outlive their reasons.

Running, where the design has to be held

Running is where the provider side genuinely carries most of the weight: delivery, timing, capture, and scoring all automate well. The client work hides at the edges. Structure has to be held in the room, session after session, against exactly the erosion the decay section opened with. The integrity posture has to stay proportionate to the stakes and announced to the person being assessed, the settlement the briefing on assessment security and governance argues from both directions. And the sitting has to remain something a serious candidate will finish and respect, the standard the briefing on candidate experience holds the whole funnel to. Each of these is a policy the platform can enforce only after the organization has chosen it.

Reading, with the controls on

Reading is the quadrant with the strongest evidence behind it and the weakest habits in front of it. The default should be combination by rule, scores entering the decision through a mechanism fixed in advance, the case the briefing on mechanical versus holistic judgment makes at length and the meta-analytic record backs (Kuncel, Klieger, Connelly, & Ones, 2013). Around the rule sit the controls: the band printed beside any score that sits near a line, and the exception log, where overrides and integrity flags go to be reviewed by someone with the standing to reverse them, the review model the briefing on integrity flags and false positives builds. A read quadrant in good order is recognizable from one seat away: nobody in it is averaging impressions.

Learn, or the wheel stops

Learning is the quadrant that keeps the others honest, and the one that exists least often, because everything it needs sits on the far side of the hire, where assessment teams rarely have standing. Its artifacts: quality of hire, recorded against a definition fixed in advance, the subject of the next section; early leavers, logged against what selection could have seen, the measurement the briefing on early attrition specifies; profiles that keep working after the offer, inside the fences the briefing on assessment data after hiring builds around them; and the estimates themselves, updated as local outcomes accumulate, the continuous-validation machinery the local-validation playbook owns and this capstone only points at. All of it presumes a platform where the score and the later outcome sit in one record, and all of it still waits, on most such platforms, for an organization that decided to look.

The arrow back from learn to define is the segment most programs never install. Learning that does not revise a definition is a report, and reports circulate without changing what the next requisition selects for. When the return is wired, artifacts start to move: a criterion definition gets rewritten because the outcome record showed the old one was scoring tenure and calling it performance; a cut policy shifts because the applicant pool that justified it no longer exists; a role profile loses a demand nobody was ever able to observe. Those revisions are the only proof that the model is running rather than merely drawn. A program that has never changed a definition it wrote at launch is not stable. It is asleep.

What should impress about Figure 1 is how little of it is measurement. Instruments, delivery, integrity capture, reporting, analytics: the material half can arrive as a service, and every quadrant still turns on an act no service performs: a definition chosen, a structure held, an exception reviewed, an outcome recorded. The hub of the wheel is drawn as a cadence ring for that reason: the quadrants are what the program does, and the hub is why it keeps doing it.

Quality of hire is written before it is measured

The learn quadrant depends on a measurement most organizations have never specified. Quality of hire circulates as a metric to be adopted, something a dashboard would supply if only the right fields were mapped. It behaves like a definition to be written, and the writing is client work, because the outcomes it names live in systems the assessment team neither owns nor sees.

This is the criterion problem, returning at program scale. Every predictive claim in this series, beginning with the flagship account of what predicts job performance, is a claim about a criterion somebody recorded, and the synthesis that anchors the series judges selection methods against outcomes an organization bothered to capture (Schmidt & Hunter, 1998). An organization with no recorded outcome still makes predictions. It simply cannot be wrong about them, and every instrument in the funnel keeps whatever reputation it arrived with.

Unpacked, the term resolves into branches that answer different questions and age at different speeds, and Figure 2 pairs each with the briefing that carries its evidence. Retention is the branch most programs already hold, because leavers get recorded whether anyone planned it or not, and the briefing on early attrition is largely an argument about how much selection-relevant signal sits unread there. Process integrity is the branch most programs omit, and it governs whether the others mean anything: a performance figure attached to an unverified sitting measures something, and nobody can say what.

The temptation here is a single index: branches weighted together into one score a leadership team can watch move. Resist it for the reason the branches exist. They disagree, and the disagreement is the information. A cohort that performs early and leaves inside the year has said something precise about what the definitions selected for, where an index would report a middling figure and no direction. Keep the branches apart, keep them dated, and let the composite live in a sentence.

One rule protects the whole exercise: fix the definition before the campaign it will judge. A criterion chosen after results are in will confirm whatever the program already believed, because some outcome always exists under which the last cohort looks fine. Written first, the definition becomes a commitment the learn quadrant can hold the define quadrant to, and that commitment is what gives the wheel's return arrow something to carry.

Nothing recurs unless somebody chairs it

A wheel with no clock is a diagram. The clock this model runs on has a quarterly hand, an annual hand, and an interrupt, and each answers a different question about how fast its subject matter goes stale.

The quarterly read is short and mostly about noticing. Funnel health since the last one. The exceptions logged and what became of them. Anything that moved without being asked to: pass rates drifting at a stage, a category of integrity flag that changed volume for no stated reason. Its output is a short list of things to examine properly, and its discipline is that the meeting happens on the date even when nothing appears to be wrong. A read that convenes only when somebody is worried has become an incident review with a recurring slot.

The annual deep pass opens the artifacts themselves. Are these still the right instruments for these roles, and does the evidence behind each still read as it did when it was filed? Do the cut policies still carry their rationales, and do those rationales still describe the applicant pool? Are the norms still current? The testing standards treat this as ordinary practice: a test's intended uses get documented, and the documentation gets revisited periodically (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014). The pass is annual because these artifacts decay on the timescale of hiring markets, which is slower than a quarter and faster than a memory.

Then the interrupts, the part organizations most reliably omit, because they are conditions and the rest of the calendar is dates. A new region opens, and what feels like extending the program is closer to launching instruments into a population that has never taken them, the argument the briefing on global assessment programs makes at operating depth. A new role family arrives with no role profile, no criterion definition, and no cut rationale to inherit. Hiring volume moves far enough that the binding constraint changes identity, the shift the briefing on high-volume hourly hiring is built around. Each should trigger a pass through the define quadrant. Each usually triggers a requisition.

Ownership is where operating models get misread as headcount requests, because the instinctive answer to "who owns this" is a number of people. This model asks for chairs. A program owner, accountable for the model as a model and seated where the quadrants are all visible at once. An assessment owner per role family, holding that family's definitions and answering for their currency. An integrity reviewer, with the standing to overturn a flag and enough independence to be worth asking. Each is a named seat with a recurring date, and each can be held by someone who already has a job. Attention on a schedule is a smaller ask than a function, and attention is what the model is short of.

Across brands, regions, or business units the chairs multiply sideways: one program owner still, and more assessment owners, because role families are where definitions live. Enterprise programs that skip that end up with a central team defining roles it has never hired for, and local hiring routing around the definitions.

Most programs are instrumented without being operated

Programs sit somewhere along a gradient of posture, and most serious organizations sit in its middle, which is a flattering place to be stuck. The instruments are good. The process is structured. The reports get read. Nothing in the quarter looks wrong, and nothing in the quarter changed a definition, because the return arrow from learn to define was never wired to anything and no seat exists whose job is to pull it.

What separates the postures in Figure 3 is dateability. In the operating posture, somebody can name a definition that changed, say when it changed, and produce the outcome record that caused the change. That question does more diagnostic work than any maturity questionnaire, and it can be answered in a meeting.

The instrumented posture is where evidence-based hiring most often stalls, and it is stable because its failure is invisible from inside. Scores keep arriving, decisions keep getting made, and the artifacts age at a rate no dashboard displays. Organizations leave it the way anyone leaves a comfortable local optimum: one recurring item on a calendar, with a chair attached.

Reaching the operating posture changes what a program can say about itself. Asked why a cut line sits where it sits, it answers with a date and a reason. Asked whether the assessment is working, it answers with the branches in Figure 2 and the direction each is moving. Asked what it learned last year, it names a definition it rewrote. Those answers are small, and none can be produced without having run the model for at least one turn of the wheel, which is why the postures advance at the calendar's pace.

The ad hoc posture is where every program starts, and its next move is the smallest of the set: writing down what good performance will mean for one open role requires no budget, no software, and no permission. It also has the highest yield, because without a written criterion almost nothing the rest of this series established can be applied at all.

Name an owner, then give the model a calendar

The starting instructions for a capstone should fit inside a quarter, and these do. Name the program owner first, because every other instruction needs a seat to fail in, and give that person the standing to change a definition, which is a different grant from the standing to report on one. Put the quadrants on the calendar next, with dates: a quarterly read with an agenda, an annual deep pass with the artifacts named in advance, and a written trigger list, so that a new region or a new role family cannot arrive unannounced.

Define quality of hire before the next campaign opens, and write the definition where the campaign can see it. Sophistication is optional; existing in advance and naming which outcomes will count is not. Attach the recording at once, because whatever system holds performance, ramp, and leavers has to hand those back to the assessment record. A definition nobody can compute gets renegotiated the first time it becomes inconvenient.

The last instruction splits the ledger deliberately. Hold the provider half to its part: instruments with evidence behind them, delivery that behaves, capture that stands up, reporting that shows its working. Keep the client half in-house and on the calendar, because it decays without a chair and no supplier can hold it for you. Participant data is the duty the service-delivery standard writes into both halves at once: the obligations attached to the data an assessment generates sit with the client organization as well as with the provider that collected it (International Organization for Standardization, 2020a, 2020b). Splitting the ledger divides the labor. It does not move that duty off your side of it. An organization running both halves will still make bad hires. It will know, and it will know why, and its definitions will move because of it.

Sixty briefings, and one order in which their arguments have to be kept. The quadrants are that order.

Every briefing in the series placed the same wager under a different subject heading, and each left the placing to somebody with a calendar. That is where it sits now: in the quarter you schedule and the definition you rewrite. The wager was always evidence over intuition. A calendar is what keeps the argument.

Where 5Profiler stands

The provider half of the arrangement this article describes is what 5Profiler supplies: instruments that arrive validated and documented, delivery and integrity capture that run as configured, and reporting that shows its working rather than asserting it. The client half of that arrangement remains yours, unchanged by anything we ship. The two meet at the record, which is where the analytics layer sits: assessment scores and the outcomes that arrive later are kept in one place, so a quality-of-hire definition has somewhere to be computed instead of argued about, and a role profile can be revised against what the outcomes showed instead of against what anyone remembers.

Read the science behind the platform · See it on your roles

References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
  2. International Organization for Standardization. (2020a). Assessment service delivery — Procedures and methods to assess people in work and organizational settings — Part 1: Requirements for the client (ISO 10667-1:2020). International Organization for Standardization.
  3. International Organization for Standardization. (2020b). Assessment service delivery — Procedures and methods to assess people in work and organizational settings — Part 2: Requirements for service providers (ISO 10667-2:2020). International Organization for Standardization.
  4. Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98(6), 1060–1072.
  5. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.
  6. Society for Industrial and Organizational Psychology. (2018). Principles for the validation and use of personnel selection procedures (5th ed.). Industrial and Organizational Psychology, 11(Suppl. 1), 1–97.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.