A candidate pops balloons for 12 minutes. Each balloon can be pumped larger for a bigger banked reward or cashed in before it bursts, and when the session ends an algorithm turns the clicks into a number. The number arrives formatted like a test score, sits in the applicant tracking system like one, and gets compared across candidates like one. Everything depends on a question the animation cannot answer: is it one?
That question is the entire business case for game-based assessments. The pitch has two parts: a measurement claim, that the game captures the construct an established cognitive test captures, and an experience claim, that candidates enjoy the capture enough to think better of the employer for it. Neither claim can be verified by playing the game. Both are empirical, and both now have evidence worth reading before any contract is signed.
The pitch found its moment. Games generate dense telemetry, a stream of timestamped decisions where a multiple-choice item yields one response, and machine scoring made that telemetry usable. At the entry level, with its huge applicant volumes and its employer-branding anxieties, an assessment that candidates might describe as fun sounded like a competitive weapon. That history explains why the measurement question is now urgent. It does not answer it.
The evidence covers the whole claim structure. A 2024 meta-analysis pools every published estimate of how closely game-derived scores track traditional cognitive measures. A controlled comparison tests the experience claim directly, holding the abilities constant while the wrapping varies. And one deployment, documented end to end in an open-access journal, shows what the category looks like when the full psychometric record is made public.
The exhibits agree on a shape. Game-derived scores carry genuine cognitive signal. They sit nowhere near interchangeability with the tests they echo, the games engineered for measurement do no better than the games that were not, and the engagement advantage fails to appear in controlled comparisons built to isolate it.
Game-based assessments converge with the tests they would replace, partway
Bipp, Wee, Walczok, and Hansal (2024) assembled the meta-analytic answer to the measurement claim: 807 effect sizes from 52 independent samples in 44 papers, 6,139 test takers in all, published in the Journal of Intelligence. The quantity they estimated is convergent validity, the correlation between a new instrument's scores and an established measure of the construct it claims to capture. For a product category whose promise is the familiar construct in a friendlier wrapper, convergence is the number on which everything else rests.
The pooled result is r = .30 as observed, with a 95% confidence interval running .26 to .34, rising to .45 after correction for unreliability, an estimate of how the two score families would correlate if neither carried measurement error. Figure 1 plots both readings, and both deserve to survive the sales meeting. An observed .30 whose interval sits comfortably clear of zero certifies that the games touch real cognitive variance; whatever the balloons are eliciting, part of it is the intended ability.
The corrected figure is the ceiling, and the ceiling is instructive. A correlation of .45 means the two score families share about a fifth of their variance (.45 squared, computed: roughly 20 percent). The rest is everything the correlation cannot see: measurement error, plus whatever a game demands that a test does not. Related, not interchangeable.
Which of the two numbers should govern a decision depends on what the decision is about. The corrected .45 describes constructs: the abilities the games reach overlap meaningfully with the abilities established tests reach, and a research program can build on that. The observed .30 describes instruments: the scores an organization would actually put into a ranking carry their unreliability with them, and every ranking is built from scores. Procurement lives at the observed number.
Whether cognitive ability is worth measuring in the first place is a settled argument with an unsettled public reputation, and it is the subject of the general-mental-ability briefing. This briefing prints no validities against job performance, because the meta-analysis in view estimates none. What it estimates is distance: how far the game score sits from the score it is named for.
The distance matters operationally. Two instruments correlated at .30 rank an applicant pool in visibly different orders, so an organization that swaps a validated cognitive test for a single game has not changed the delivery of a fixed decision; it has changed who gets hired. The corrected .45 softens that statement without retiring it. Even at the ceiling, the game and the test disagree about most of what they see.
The single game is the weak unit of measurement
Inside the same dataset, the moderator analyses carry the practical instruction. Convergence rises with aggregation: single-game scores correlated .29 with traditional cognitive measures, while composites built from several games reached .38. The mirror image holds on the reference side, where convergence with a single cognitive measure ran .28, against .39 when multiple cognitive measures were combined (Bipp et al., 2024).
The pattern predates the format by decades, restated here in telemetry: composites beat single indicators, because aggregation averages away error and idiosyncrasy. The flagship briefing on what predicts job performance meets that logic in every method it reviews, and it does not stop applying because the items are animated. A balloon game is an item type. No psychologist would score a cognitive test on its three most colorful items and call the measurement finished.
The aggregation instruction collides with the category's own sales pitch. The format is marketed as brief and painless, and a composite of several games is neither as brief nor as simple to assemble as one game with one score. The trade is real: a single short game maximizes the experience pitch, a battery of several games maximizes convergence, and the moderator estimates quantify what each configuration gives up. An organization can choose either, but it should choose knowingly, with .29 and .38 in view.
Reliability sets a second limit from underneath. Across the meta-analytic dataset, game measures averaged a reliability near .71, and reliability caps what any single score can carry: at that level, each score arrives with an uncertainty band wide enough that single-game verdicts overreach. The measurement-error briefing works through what bands of that width do to decisions; here it is enough to say that the .71 belongs in any conversation that quotes the .45.
Purpose-built games measured no better than repurposed ones
The meta-analysis contains a moderator result that should reorganize how these products are evaluated. Some games in the literature were designed from the outset to measure cognitive ability. Others began as entertainment and were repurposed as assessments afterward. The purpose-built group did not converge more strongly with traditional measures than the repurposed group (Bipp et al., 2024); Figure 2 shows the two at parity because parity is the finding.
The null reads like a technicality and functions like a policy. A vendor's account of its design process, the psychologists consulted, the target construct, the intended mapping from mechanics to ability, is an account of intentions. Design intent is not construct validity, and the first cannot be audited into the second. What separates a measuring game from an entertaining one is evidence gathered after the game exists.
The null admits a deflationary reading and a cautionary one, and they converge for an evaluator. The deflationary reading: cognitive load is easy to generate, and an entertainment game full of timed decisions, spatial tracking, and memory demands can tax the intended abilities without anyone having planned it. The cautionary reading: deliberate construct mapping is harder than it looks, and purposeful design has not yet produced a measurable convergence advantage over accident. On either reading, the label on the box says nothing about the score, and only evidence says anything at all.
Landers and Sanchez (2022) supplied the category's vocabulary, distinguishing game-based assessments from gamified assessments and gamefully designed assessments, and attached the conclusion that travels with the taxonomy: the validation obligations do not soften with the format. Telemetry is a set of item responses. It requires the same evidence chain as any other item type, from scoring rationale to reliability to convergence to fairness, and the chain has no shortcut for charm.
The obligation is shared with every algorithmically scored instrument now arriving in hiring. A model that scores interview recordings makes the parallel promise in a different medium, familiar construct through a novel pipeline, and the briefing on AI-scored interviews traces what auditing that promise involves. The standards transfer between the two cases without modification, which is the point of having standards.
One deployment shows what the category can document
Leutner, Codreanu, Brink, and Bitsakis (2023) published, in Frontiers in Psychology, the account this article would otherwise have had to hypothesize: a machine-learning-scored game-based assessment of cognitive ability, deployed in live recruitment, with its psychometrics reported in the open. The platform goes unnamed here, as vendors do throughout this series; the paper is the exhibit, and the paper is unusually complete.
Three results anchor it, and Figure 3 gives each its own scale. Against traditional cognitive measures, the game's scores reached a concurrent validity of r = .50, with a 95% confidence interval of .43 to .56 (concurrent meaning the game and the reference measures were taken close together in time). Retested after roughly four months, scores correlated .68 with their first administration (N = 102), which speaks to the applicant who plays again in a new hiring cycle: the number that comes back is recognizably continuous with the first. And across 4,778 applicants, the deployment posted a Net Promoter Score of 58, a satisfaction index on which positive values mean recommenders outnumber detractors.
The paper also examined fairness, in a sample of 3,107 applicants, and reported adverse-impact ratios above the conventional threshold across the groups examined. This briefing leaves the result qualitative by design; the ratio and the threshold belong to the adverse-impact briefing, which explains what such ratios do and do not establish. The instructive part is the disclosure itself: fairness evidence, computed on real applicants, published where a skeptic can read it.
Publication has a second function beyond the numbers: it fixes the claims. Psychometrics that exist only in a slide deck can be revised for each audience; psychometrics in a peer-reviewed journal are commitments, with methods a rival or a regulator can inspect and a record the vendor cannot quietly amend. A published fairness analysis, whatever its result, also sets a floor for the category, because the next vendor can now be asked why its own file is thinner.
Which product achieved these numbers matters less than the demonstration: nothing about the format prevents a vendor from documenting convergence, stability, and fairness to the standard the rest of testing accepts. A concurrent .50 also sits above the pooled observed .30 from the meta-analysis, a reminder that the category average and a well-built instrument are different objects. The average describes the market. The file describes the product in front of you, and only the file should be allowed to close the sale.
The engagement premium is contested where it has been tested
The experience claim rests on firmer ground than skeptics grant and thinner ground than the marketing implies. On the construction side, Georgiou, Gouras, and Nikolaou (2019) developed a gamified situational-judgment assessment for selection and validated it, settling a threshold question: gamification in hiring can be engineered to psychometric standards rather than sprinkled on afterward.
Whether candidates reliably prefer games is the part the evidence contests. Ohlms, Melchers, and Kanning (2024) ran the comparison directly and found construct-related validity evidence for a game-based cognitive measure, alongside applicant reactions that were not superior to reactions to conventional tests measuring the same abilities. The design carries the weight here: with the abilities held constant and only the wrapping varying, the advertised reactions advantage did not appear.
The two findings about experience are compatible. A Net Promoter Score of 58 describes one deployment, with its particular applicant pool, communication, and production polish (Leutner et al., 2023); the controlled comparison isolates the format itself and finds no general premium (Ohlms et al., 2024). Satisfaction in a deployment is evidence that a specific implementation landed well. It licenses nothing about the category.
There is also a reason the premium may fail to generalize that has nothing to do with design quality. An applicant under real stakes reads whatever the employer sends as a test, because it is one; a candidate whose next job depends on the outcome experiences a balloon as an item, and the pastel palette does not change what is at risk. On that view, the surprise in the reactions data dissolves. Candidates are serious people making a serious decision, and they respond to the substance of the process more than its surface.
What moves applicant reactions, measured across formats, is perceived job-relatedness and being treated with respect, and the candidate-experience briefing assembles that evidence. An organization whose goal is a shorter, less tedious sitting also has a route that never leaves validated instruments: adaptive testing, which trims the sitting by choosing items instead of adding animation.
Buy the evidence file, not the art style
The evidence supports a specific posture: game-based hiring tests are usable where their psychometrics are documented, and undecidable where they are not. Procurement is the moment that distinction can be enforced, because it is the moment the organization can still ask for the file. For any game-based assessment under consideration, the file has four sections worth demanding.
- Convergence. The correlation between the game's scores and an established measure of the construct it claims, in a sample the vendor can describe. The pooled .30 observed and .45 corrected (Bipp et al., 2024) are the reference points against which a vendor's number can be read.
- Reliability, score by score. Reported for each score a recruiter will see, and read against the meta-analytic mean of about .71 for the category (Bipp et al., 2024).
- Retest stability. Estimated over a window resembling a real re-application cycle; the published benchmark is .68 over roughly four months (Leutner et al., 2023).
- Fairness documentation. Computed on applicants like the organization's own, and stated plainly enough to survive an adverse-impact review.
The moderator evidence adds a configuration rule: weight composites over single games. The gap between .29 and .38 recurs on the reference side, which marks it as the aggregation logic at work everywhere in selection, showing up here in game form. A vendor whose product reports one score from one game is asking a single item type to do a battery's job.
Placement in the funnel follows from the psychometrics. A single game carrying .29 convergence and the category's average reliability is defensible as one signal inside a broader composite and hard to defend as a sole gate. Aggregated with validated instruments, its noise gets absorbed; installed as the only hurdle, its noise decides careers. The narrow claim the evidence sustains, real but partial cognitive signal, sets exactly how much decision weight the score can claim.
Reactions deserve a local pilot, because the controlled evidence says the premium cannot be presumed while the deployment evidence says a well-run implementation can post strong satisfaction numbers. Running the game and the incumbent process on comparable applicant samples, and measuring reactions in both, replaces an assumption with a finding at the cost of one hiring season's patience. A pilot also surfaces the operational facts no journal article can supply: how the game behaves on the organization's own devices and networks, how long candidates actually spend in it, and how its scores line up against the instruments already in use. If the premium is real for a given population, the pilot will show it; if it is not, the organization has learned that before the format hardened into policy.
The stakes concentrate where the pitch is loudest: high-volume campus hiring, where thousands of first-time applicants, an employer-branding motive, and compressed timelines make engagement the easiest thing to sell and the psychometrics the easiest thing to defer. Volume cuts the other way. Across thousands of decisions a season, the difference between an aggregate and a lone game stops being a rounding concern, and so does the category's average reliability. A graduate program that would never adopt an interview format without structure should not adopt a scoring instrument without a technical file, however good the trailer looks.
The Standards for Educational and Psychological Testing apply with their full weight regardless of delivery format (American Educational Research Association et al., 2014): a game that produces scores used for selection decisions is a test, carrying a test's obligations for evidence, documentation, and use. The practical corollary is that the documentation an organization already expects from an assessment platform, technical manuals, reliability tables, validity evidence, is the documentation a game vendor owes, unchanged. One question this briefing deliberately leaves open is whether game formats are harder or easier for candidates to outsource to software; the studies reviewed here measure nothing about that, and the question belongs to assessment design in the AI era.
The games will keep improving, and this literature, read plainly, argues for holding the door open rather than closing it: the format has produced at least one instrument with a documented claim to serious measurement, and the route it took is repeatable. The route runs through evidence, and organizations control the incentive by what they agree to accept. A game is a delivery mechanism. The psychometrics have to be bought separately, and the standards for them never changed.
Where 5Profiler stands
5Profiler settles the measurement question before the medium question. The assessments candidates sit are validated instruments with the science published and inspectable, and the candidate experience is engineered through clarity: plain instructions, visible progress, explained purpose, rather than gamification layered over the construct. Where a shorter sitting matters, adaptive testing provides it. So when a session ends and a number comes out, the number has answers ready for the questions this briefing opened on: what it measured, how reliably, and on what evidence, before anyone asks how it felt to produce.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
- Bipp, T., Wee, S., Walczok, M., & Hansal, L. (2024). The relationship between game-related assessment and traditional measures of cognitive ability—A meta-analysis. Journal of Intelligence, 12(12), 129.
- Georgiou, K., Gouras, A., & Nikolaou, I. (2019). Gamification in employee selection: The development of a gamified assessment. International Journal of Selection and Assessment, 27(2), 91–103.
- Landers, R. N., & Sanchez, D. R. (2022). Game-based, gamified, and gamefully designed assessments for employee selection: Definitions, distinctions, design, and validation. International Journal of Selection and Assessment, 30(1), 1–13.
- Leutner, F., Codreanu, S.-C., Brink, S., & Bitsakis, T. (2023). Game based assessments of cognitive ability in recruitment: Validity, fairness and test-taking experience. Frontiers in Psychology, 13, 942662.
- Ohlms, M. L., Melchers, K. G., & Kanning, U. P. (2024). Can we playfully measure cognitive ability? Construct-related validity and applicant reactions. International Journal of Selection and Assessment, 32(1).