Evidence, published.
Long-form research briefings on the science of talent decisions — what the peer-reviewed literature actually shows, with sources you can hand to your assessment and legal teams.
Read the evidence behind the platform
Every briefing is built from published, citable research — meta-analyses, standards documents and regulatory texts — and closes with what the evidence means for how organizations assess people.
What actually predicts job performance
Meta-analyses spanning a hundred years of selection research agree on something uncomfortable: the methods organizations trust most predict least. The evidence, and the 2022 re-analysis that reordered it.
Personality scienceThe Big Five is not a personality quiz
Type indicators reclassify a large share of retakers within weeks. Trait models replicate across cultures and decades — and predict work outcomes. Why the distinction matters more in hiring than anywhere else.
Personality scienceFive scores hide more than they reveal
Two candidates can post identical Big Five profiles and behave in opposite ways on the job. The bandwidth–fidelity debate, the meta-analytic verdict, and the case for measuring at facet resolution.
MeasurementThe best assessments are shorter
Computerized adaptive testing reaches fixed-form precision with roughly half the items, holds precision constant across the ability range, and starves cheating economies of a stable answer key. The psychometrics, explained.
Economics of hiringThe money math of better hiring
A formula published in 1949 converts assessment validity into currency — and it has survived every honest attack on its assumptions. The math, the caveats, and why selection quality is still systematically underpriced.
Cost of getting it wrongThe true cost of a bad hire
Direct replacement runs about a fifth of salary; fully loaded estimates run to twice it. What a mis-hire really costs, where the money hides, and the selection arithmetic that prices the risk before the offer.
Cost of getting it wrongThe interview illusion
Free-form interviews predict performance at .19 — and in controlled studies they made predictions worse than no interview at all. Why the most trusted method in hiring earns the least trust, and what fixes it.
Cost of getting it wrongRésumés are claims, not evidence
Years of education predict performance at .10, experience at .18 — and identical résumés draw 50% more callbacks with a different name at the top. The case for screening on measurement instead of paper.
Cost of getting it wrongNinety-day failures
Person–job fit correlates −.46 with the intention to quit — and it is measurable before the offer. Why new hires leave early, which signals predict it, and what each preventable exit is worth.
Personality scienceWhat conscientiousness actually buys
It is the trait that generalizes across nearly every job studied — and the enthusiasm around it has outrun the evidence. What it predicts, which facets carry the weight, and why grit is largely the same construct rebranded.
Personality scienceThe cost of volatility
Neuroticism is the strongest personality correlate of job dissatisfaction, and dissatisfaction is where turnover and counterproductive behavior begin. What emotional stability predicts, and where it matters most.
Personality scienceCan candidates fake a personality test?
Applicants do score higher than incumbents, and told to fake, people can. Yet validity survives more of it than the fear suggests. What faking does to your rankings — and the countermeasures that hold up.
Personality scienceInterests are not competence
Interest inventories are the most popular instruments in career guidance and among the weakest in selection. They predict what people choose and stay with far better than how well they do it.
Personality scienceThe dark side of strengths
Most trait–performance relationships are curves, not lines, but hiring treats them as lines and rewards the extremes. Where more stops being better, and how the strength that got someone promoted becomes the flaw that ends it.
Fairness & complianceThe four-fifths rule
A one-line arithmetic test from 1978 still decides when a selection procedure needs defending. How adverse impact is measured, what the diversity-validity research actually says, and the design choices that move both numbers.
Fairness & complianceStandardization is a fairness technology
Bias enters selection through discretion: shifting criteria, unequal questions, evidence weighed after the preference forms. The research case that structure closes those doors — and training alone does not.
Fairness & complianceThe rules have arrived
Hiring algorithms are now regulated technology on three continents: bias audits in New York, high-risk obligations in Europe, consent architecture in India. What the rules require, and what they quietly reward.
Fairness & complianceCandidate data is a liability until governed
Every assessment funnel accumulates sensitive records of people who will never be employees. The privacy principles that govern that archive, who is responsible for what, and the schedule most programs never set.
Fairness & complianceThe assessment is the employer brand
For most applicants, the assessment is the highest-resolution contact they will ever have with your company. What the applicant-reactions research says drives their judgment — and what that judgment goes on to cost.
Integrity & proctoringTesting without a proctor
Unproctored testing is a trade, not a failure: reach and speed against certainty about who did the work. What the research says actually happens when the proctor leaves, and the two-stage designs that keep the trade honest.
Integrity & proctoringDeterrence beats detection
The best integrity event is the attempt that never happens. What deterrence research says about visible monitoring, why certainty outworks severity, and how to protect a funnel without treating every candidate as a suspect.
Integrity & proctoringAssessment in the age of ChatGPT
A frontier model scores near the top of a bar exam its predecessor failed. The formats that collapse under that capability, the designs that survive it, and why betting on detection is the wrong response.
Integrity & proctoringThe flag is not a verdict
Run a highly accurate detector across thousands of honest candidates and it will still accuse the innocent at scale. The base-rate arithmetic of integrity flags, and the review process that keeps them defensible.
Integrity & proctoringSecurity is a program, not a feature
Locks are bought; security is run. What the test-security guidelines actually prescribe — risk assessment, prevention, detection, response and review — and the vendor-employer ownership split most contracts never write down.
Integrity & proctoringWho actually took your test?
In recent samples, nearly 16% of students admit to paying someone to do their work — and a 10% cheating rate can capture a third of your top-decile slots. The fraud vectors in remote hiring, and the integrity stack that closes them.
Applied playbooksCampus season, by the numbers
On campus every file looks the same, and the season gives you one pass. How validated batteries, structured late rounds and season telemetry turn volume from a threat into a signal.
Applied playbooksHow to actually run a structured interview
The evidence for structured interviews is settled; the practice is not. The field manual: job-derived questions, behaviorally anchored scales, independent ratings and rule-based scoring.
Applied playbooksCoding assessments that predict
A coding task only measures when it is built like an instrument: standardized environment, bounded scope, dimension rubrics and integrity by construction. The design manual.
Applied playbooksThe file that dies on its best day
The funnel’s most expensive artifact usually goes unread after the offer. How profiles brief the first one-on-one, shape onboarding and seed development — as tools, never verdicts.
Applied playbooksThe standards-based assessment program
The published standards of the field reduce to six commitments an ordinary organization can adopt — each with a standard behind it, a briefing in this series, and an artifact a working program produces.
Ability & skillsThe most contested number in hiring
For forty years the field said cognitive ability predicted performance at .51. In 2022 the estimate fell to .31 — not because ability stopped mattering, but because the correction arithmetic changed. What happened, and what it means for buyers.
Ability & skillsThe test that asks what you’d do
The SJT predicts performance at .26 — and measures ability or personality depending on how the question is phrased. What the evidence says about hiring’s most shape-shifting method, and how to deploy it deliberately.
Ability & skillsWatch them work
Work samples held the top of the validity table for a generation; the modern estimate is .33. And the method measures what a candidate can do today — not what they will do daily. The evidence, and the design choices that decide whether a simulation earns its cost.
Ability & skillsThe method that works for the wrong reasons
Assessment centers predict performance, and forty years of research says they do not measure the dimensions on their scorecards. Exercise effects, the general factor, and what that means for the most expensive method in selection.
Ability & skillsThe reform that moved postings, not hires
Companies tore degree requirements out of their job ads — and kept hiring graduates. The gap between announced skills-based hiring and realized hiring is a measurement gap, and it has a fix.
Personality scienceThe room-filler myth
Extraversion predicts sales at .15, and generic measures manage .10 against job performance — yet the assertive half of the trait predicts, and ambiverts out-earn strong extraverts. The evidence, split properly.
Personality scienceAgreeableness cuts both ways
In team-based work agreeableness is the strongest personality predictor (.33); across four studies it also predicts lower earnings, and in zero-sum bargaining it is a measured liability. Context decides the sign.
Personality scienceThe trait hiring wrote off
Openness predicts overall job performance at .04 — the bottom of the table — and is also psychology’s most reliable creativity marker, a top trait predictor of training success, and the trait that predicted adaptation when the rules changed.
Personality scienceOld traits, new boxes
Mixed-model EQ questionnaires are reproducible from seven constructs psychology already had (R = .79), and their validity drops to nil once those are controlled. Growth mindset moves achievement by d = 0.08. The audit.
Personality scienceSame score, different people
Clustering 1.5 million profiles finds four replicable personality types — as density peaks, not boxes. And scoring the same items at finer grain roughly doubled the variance explained. The case for reading configurations, honestly.
Fit, teams & cultureFit predicts staying, not performing
Values congruence predicts commitment at .51 and intent to stay — and job performance at .07, with a credibility interval that includes zero. How to measure fit honestly, and which decisions it should own.
Fit, teams & cultureHire for the roster, not just the role
Team-average agreeableness and conscientiousness lift performance — and one low scorer can matter more than the mean. The operators, the evidence, and the craft of team-aware hiring.
Fit, teams & cultureCulture fit is a claim, not a measurement
In elite firms, more than half of evaluators ranked fit — similarity in leisure, background, and self-presentation — as their top interview criterion. And 73% of the fit signals recruiters name are private to the recruiter.
Fit, teams & cultureThe promotion spreadsheet is wrong
Doubling a salesperson’s numbers raised promotion odds by a third — and predicted a 6.1% decline in every subordinate’s sales once they managed. Promotion is a selection decision; the evidence for treating it like one.
Fit, teams & cultureThe office was part of the instrument
Remote and hybrid work measurably work — and autonomy changes who performs: conscientiousness predicts far more steeply when nobody is watching. The evidence, and the re-weighting hiring owes a distributed role.
Scores & decisionsA score is a comparison
The same performance can read as ordinary or exceptional depending on the norm group beneath it — and norms age at a measured 2.31 points a decade. How to read any assessment number: compared with whom, measured when.
Scores & decisionsWhere to draw the line
No method discovers the “true” pass mark; the defensible ones structure the judgment, document it, and monitor its consequences. The standard-setting toolkit, and the trade every cut makes between quantity and quality.
Scores & decisionsEvery score wears a band
At reliability .90 a 95% confidence band spans about ±0.6 SD; at .70, about ±1.1 — and near a threshold the band becomes the probability the verdict is wrong. The psychometrics every report reader is owed.
Scores & decisionsThe formula beats the debrief
Seventy years of evidence, one direction: combining the same evidence by explicit rule predicts job performance at .44 against .28 for expert synthesis — a more-than-50% improvement the debrief room gives back.
Scores & decisionsWill it work here?
The intuitive answer — run your own validity study — is a noise generator below several hundred hires. Validity generalization, transportability, synthetic validity and Bayesian updating: the evidence playbook for real org sizes.
The AI eraThe interview with nobody in the room
Algorithmic interview scores lean on what candidates say, not how they look — and the medium itself lowers ratings (d = −.41). The evidence on automated video interviews, and the questions that separate instruments from demos.
The AI eraThe game is not the measure
Game scores correlate .30 with traditional cognitive tests — .45 corrected — and purpose-built games do no better than repurposed ones. What gamified assessment can honestly claim, and what buyers should demand.
The AI eraThe grader that changed overnight
On benchmark essays ChatGPT agrees with human raters inside the human range — and the same product’s accuracy on one bounded task fell from 97.6% to 2.4% in three months. The governance that makes machine grading usable.
The AI eraThe candidate’s other window
By 2023, 17% of surveyed early-career candidates had already used generative AI in applications and 70% expected to — while GPT-4 outscored 98.8% of humans on verbal reasoning. The policy spectrum, priced honestly.
The AI eraThe audit nobody posted
A year into the first mandated bias-audit regime, 391 employers yielded 18 posted audits. What transparency law actually asks, why “explain the model” is the wrong target, and the three-layer architecture that works.
Applied playbooksVolume is a constraint, not an excuse
Frontline funnels screen thousands on phones between shifts. The evidence says what survives the device, what shortening really costs, and why the criterion that pays is ninety-day survival — not speed-to-fill alone.
Applied playbooksHire the pipeline, not the party trick
The sales meta-analysis splits every predictor in two: cognitive ability predicts manager ratings at .40 and objective sales at .04, while conscientiousness predicts the register itself at .41. The playbook follows the split.
Applied playbooksWhen one hire is 10% of the company
Below fifty employees every hire is a measurable share of the firm — and hiring is at its most informal exactly where each decision weighs most. The minimal structure that keeps decision quality before you can afford more.
Applied playbooksScores need passports
One battery across three continents only stays one measurement if four portability problems are handled deliberately: adaptation, measurement invariance, response styles, and the norm each region is read against.
Applied playbooksThe evidence-based hiring operating model
The field’s own standard takes two documents to describe one assessment: the provider’s half and the client’s. This capstone draws the client half — the operating rhythm that keeps evidence-based hiring true after the purchase.
Sixty briefings, from selection-science foundations to the AI era of assessment. Start with what predicts job performance, see the first series’ design in the standards capstone, or close the loop with the operating-model capstone.
Evidence over intuition, on your next hire.
A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.