Research Personality science

Can candidates fake a personality test?

Applicants do score higher than incumbents, and told to fake, people can. Yet validity survives more of it than the fear suggests. What faking does to your rankings — and the countermeasures that hold up.

Two findings about faking personality tests have sat side by side for three decades, and the assessment market has spent most of that period arguing about which one to believe. Candidates can inflate their scores, and real applicants do, in the direction the job advertisement points. Yet the statistic buyers care most about, the relationship between test scores and later job performance, has proved stubbornly resistant to that inflation, and statistically removing the supposed contaminant does not improve prediction. The inflation is real and the damage refuses to show up where everyone has been looking for it.

Buyers feel the first half of that pair long before they hear the second. The objection surfaces in nearly every procurement conversation about personality assessment, usually in the same words: what stops a candidate from simply giving us the answers we want to hear? It is not a naive question. It is the correct instinct of someone who has read a job advertisement asking for a self-starter and then read a questionnaire item asking whether you consider yourself a self-starter.

The resolution is not that one body of evidence is wrong. Validity and shortlists are different objects, computed on different parts of the same distribution: a validity coefficient is calculated across an entire pool of candidates, while hiring consumes only the top of a ranking. On a simple model set out below, with 30% of applicants inflating by 0.45 standard deviations, that minority leaves the pool's average almost untouched while taking nearly half the places on the shortlist. That is where the damage from applicant faking actually lands, and it is a different problem, with different countermeasures, from the one the objection usually describes.

What candidates can do when you ask them to

Start with the upper limit, because it is the part nobody disputes. Directed-faking studies hand respondents the instruction the objection assumes they give themselves: answer as though you badly want this job. Viswesvaran and Ones (1999), pooling that literature in Educational and Psychological Measurement, found that instructed respondents raise their scores substantially, on the order of half a standard deviation across dimensions and more on some. Nobody is confused by the items. Told what to look like, people can look like it.

That result deserves to be stated without hedging, because the defensive reflex in the assessment industry has been to bury it. A questionnaire that asks whether you finish what you start is transparent by construction, and pretending otherwise insults a buyer who can read the items as easily as a candidate can. Self-report personality measures are fakable, and the argument that follows has to survive that concession rather than avoid it.

What a directed-faking design measures, though, is a capacity rather than a behavior. The instruction strips away every friction that operates in a live sitting. A real applicant does not know with certainty which traits this employer prizes, how the scoring works, or whether a claim will be probed in an interview later, and nobody has told them to distort. The laboratory removes all of that and reports what happens when it is gone.

Reading a directed-faking effect as a prevalence estimate is a persistent error in the debate about faking personality tests, and it is made in both directions. Vendors cite the artificiality to dismiss the finding. Skeptics cite the magnitude as though it described a typical applicant. The estimate is a ceiling: it tells you the size of the lever, not how often or how hard anyone pulls it. For a buyer, the useful question is therefore narrower than "can it be faked," which has been answered. The question is how far real applicants actually move, on which traits, and with what consequence for the ranking that gets acted on.

Applicants inflate exactly where the advertisement points

Birkeland, Manson, Kisamore, Brannick, and Smith (2006) answered the first two parts of that question directly. Their meta-analytic investigation of job applicant faking, published in the International Journal of Selection and Assessment, compared people answering personality measures as actual job applicants against people answering the same instruments outside a hiring context. Applicants scored higher. The interesting part is that they did not score uniformly higher.

The differences concentrated on two traits. Conscientiousness showed a gap of about .45 standard deviations, and emotional stability about .44. Differences on extraversion, openness, and agreeableness were smaller. Figure 1 shows the shape of the finding, with the smaller three represented as the source represents them, qualitatively, because no pooled value for them belongs in this article.

Read the pattern before the magnitudes. The two traits carrying the largest gaps are the two that job advertisements request most reliably: reliability and follow-through on one side, composure under pressure on the other. Openness and agreeableness are not less admirable human qualities. They are less consistently demanded in the posting, and the applicant gaps track demand rather than desirability.

This is why "faking" is a slightly clumsy word for what the data show. Blanket lying would raise every scale, since a respondent simply trying to look good has no reason to be selective. What these applicants did looks instead like impression management aimed at a target, the ordinary human behavior of presenting the version of yourself a specific situation rewards. Social desirability in hiring is not free-floating vanity, it is job-directed.

That reframing also explains why the effect resists design solutions. An applicant who reads the advertisement, infers what the role demands, and then answers a questionnaire in that light has done something closer to self-presentation than to deception, and no wording of the items eliminates it. Self-report in a high-stakes setting is always a conversation with an audience, and candidates know precisely who the audience is. The response distortion literature is describing a feature of the situation, not a population of liars.

The contaminant that was not there

Here the evidence turns against the intuition it has just supported. If inflation contaminates scores, the contamination should be visible where it matters commercially, in the relationship between what a test says at hire and what the person does on the job. Ones, Viswesvaran, and Reiss (1996) tested exactly that in the Journal of Applied Psychology, in a synthesis whose subtitle names its conclusion: the red herring. Social desirability, the tendency to describe oneself in socially approved terms, did not attenuate the criterion-related validity of personality measures.

The second half of their result is the part that stops people. When the authors removed social desirability statistically, partialling it out of the scores, prediction of job performance did not improve. Sit with what that implies. If social desirability were noise laid on top of a true score, stripping it away should sharpen the signal, in the same way that removing a known source of measurement error usually does. It did not sharpen anything. The supposed contaminant behaved as though it were not contaminating.

Hough, Eaton, Dunnette, Kamp, and McCloy (1990) had already reported the companion finding from the applicant side, also in the Journal of Applied Psychology: intentional response distortion did not substantially degrade the criterion-related validities they observed in applicant samples. Two independent lines of evidence, one working through social desirability scales and one through deliberate distortion, arrive at the same negative result.

Why the null result holds is less settled than the result itself. One reading is that the tendency to present well is reasonably stable across people, so it shifts scores without scrambling the order candidates stand in. Another is that presenting well under stakes carries genuine job-relevant information, in which case it is not error at all. The studies cited here establish the negative finding robustly and do not adjudicate between the explanations.

The tested operation was narrow, and worth restating precisely: remove the social desirability component from a score, then check whether prediction improves. It did not. Extending that to the wider market is an inference rather than a finding, but it is a hard one to avoid, because lie scales, desirability keys, and score corrections sold as faking detection all rest on the same premise the test examined. A buyer being sold a faking-correction module is entitled to ask which evidence shows the correction improves anything a hiring team can bank.

The strongest case against self-report personality, stated in full

A reader who stops at the red herring paper will conclude the objection is dead. It is not, and the reason is that a panel of scholars said so in print. Morgeson, Campion, Dipboye, Hollenbeck, Murphy, and Schmitt (2007), writing in Personnel Psychology, argued that self-report personality measures show low validity for personnel selection and that faking remains a genuine and unresolved concern. It is a serious paper by serious people, and it deserves to be represented as they wrote it rather than as a foil.

Their case does not require refuting the null results above. It moves the argument to different ground: whether the validities being defended are large enough to carry the weight that hiring places on them. That question has only sharpened since. Sackett, Zhang, Berry, and Lievens (2022), applying less aggressive range-restriction corrections than earlier syntheses, place conscientiousness near .19 against job performance, well below the structured interview at .42. The full ranking of methods is the subject of the series' flagship briefing on what predicts job performance. A modest coefficient with a live faking concern attached is a different commercial proposition from a large one, and the panel was right to say so.

The panel's second argument is normative rather than statistical, and it is the harder one to answer. A selection instrument that applicants can move at will has a legitimacy problem even where it retains predictive value, because selection systems have to be defensible to the people they judge. A candidate who answered candidly and lost a place to someone who answered strategically has a complaint that no validity coefficient addresses. Fairness between candidates and prediction across candidates are separate goods, and an instrument can deliver the second while failing the first.

Take both bodies of evidence at face value and a genuine puzzle remains. Applicants inflate. The inflation does not register as damage to validity. Careful scholars nonetheless regard faking as unresolved. All three can hold at once, and the reason lies in what a validity coefficient describes.

Why faking personality tests damages shortlists without damaging validity

A validity coefficient answers a question about a whole pool, while a hiring process awards a small number of places at the top of a ranking and discards everything below the cutoff regardless of how faithfully the scores ordered it. The two quantities can diverge, and under applicant faking they diverge sharply.

A computed example makes the divergence concrete. Assume scores in the applicant pool are standard normal, that 70% of applicants respond without inflating, and that 30% inflate by 0.45 standard deviations, the applicant gap Birkeland and colleagues reported for conscientiousness. Nothing exotic is assumed: most people answer straight, and the inflating minority moves by an amount the published literature observed rather than the laboratory ceiling.

Under those assumptions the pool's mean shifts by only +0.14 SD. An auditor comparing average scores year over year would see nothing worth investigating. But the top-decile cutoff of the mixed pool sits near z ≈ 1.45, and above that line the composition of the pool changes character: the inflating group, 30% of applicants, occupies roughly 48% of top-decile places. Figure 2 shows the two compositions side by side, which is the comparison that matters to anyone who runs a shortlist.

This is the geometry the companion analysis of remote assessment fraud works through in detail: a normal curve thins rapidly in its tail, so a fixed head start multiplies rather than adds to the probability of clearing a high bar, and the scarcer the places, the more thoroughly a shifted minority comes to fill them. The only substitution this article makes is legal behavior for misconduct. Ordinary impression management, by candidates who have done nothing wrong and would pass any integrity check, reorders a shortlist through the same mechanism and needs no bad actors at all.

What actually raises the cost of inflating

If the damage is concentrated at the cutoff, countermeasures against faking personality tests should be judged by whether they change who ends up there. That rules out most of what gets marketed as a solution, and it points at a smaller set of design choices that reduce the payoff to strategic responding rather than trying to detect it after the fact.

Format is the lever with the clearest evidence behind it. Christiansen, Burns, and Montgomery (2005), reconsidering forced-choice item formats for applicant personality assessment in Human Performance, found that forced-choice formats reduced score inflation relative to single-stimulus formats. The mechanism is intuitive. A single-stimulus item asks whether a statement describes you, and the desirable answer is obvious. A forced-choice item asks which of several similarly attractive statements describes you best, so the strategic respondent has to choose between goods rather than between a good and a bad. The lever does not disappear, but it becomes much harder to pull in a single direction.

Composite scoring attacks the payoff from the other side. If a hiring decision rests on several instruments combined by rule, moving one self-report scale buys a smaller share of the outcome, and the effort required to move everything at once rises steeply. That is design reasoning rather than a finding, and the wider case for combining structured inputs runs through this series, including the companion piece on trait measurement and type indicators. Reading a profile against the demands of a specific role adds a further constraint by removing the reward for a generically impressive answer, since the profile a controls role requires is not the profile a business development role requires. Figure 3 sets out the ladder with its limits attached.

Two further levers belong on the ladder with weaker warrants. Warning candidates that responses are reviewed, and verifying a self-report against corroborating evidence later in the process, are both reasonable design choices, and both are commonly recommended. The evidence assembled here does not size either effect, and this article will not manufacture a number for them. A vendor who quotes a precise reduction from a warning screen should be asked for the study. Where a countermeasure does carry evidence, it belongs in the specification of the delivery platform alongside the rest of the measurement design, stated with its limits rather than as a guarantee.

What none of these measures does is eliminate the problem. Each raises the cost or lowers the return of inflating, and each leaves a self-report instrument being answered by a person with an obvious interest in the outcome. Treating that as solved is how buyers end up trusting a shortlist further than the evidence supports. The realistic target is a ranking a determined impression manager cannot climb cheaply, a lower ambition than the one usually sold and a considerably more achievable one.

Defend the cutoff, not the average

The tail argument turns faking personality tests from a philosophical worry into a procurement question, because it names exactly which part of the distribution to defend. Four rules follow.

  • Never rank on a single self-report scale. A one-dimensional ranking is the condition under which the tail capture in Figure 2 operates at full strength. Scores should enter a composite assembled by rule, with personality as one weighted input among several, in the manner the validity literature has recommended for other reasons entirely.
  • Prefer inflation-resistant formats where selectivity is highest. The concentration effect grows with the cutoff, so campus programs and high-volume funnels have the strongest case for forced-choice measurement and the least excuse for a transparent single-stimulus questionnaire.
  • Audit composition, not averages. Comparing this year's mean score against last year's will certify a reordered shortlist as unchanged. The diagnostic question is who sits above your cutoff and what independent evidence corroborates them.
  • Ask whether the validity evidence comes from applicants or incumbents. A coefficient established on employees answering without stakes describes a population that had no reason to inflate. Hough and colleagues (1990) reported their result on applicant samples, and a vendor who cannot say which kind of sample produced their numbers has not thought about the question this article is built on.

A fifth practice sits underneath the other four. Decide in advance what personality evidence is allowed to do in your process. Used as a gate, a fakable self-report scale will admit the people best at reading your advertisement. Used as one weighted input, or as a source of structured interview probes that put a claim under scrutiny with a hiring manager present, the same instrument contributes real signal at a fraction of the risk. The difference is entirely in the decision rule, not in the questionnaire.

The objection raised in the procurement conversation was right about the phenomenon and wrong about the damage. Candidates do manage their answers, and they manage them intelligently, in the direction the job advertisement points. The validity statistics survive it, exactly as the research says. What does not survive it is the assumption that the top of your ranking is populated by the people the instrument scored highest on the trait, rather than by the people who combined some of that trait with a good reading of what you asked for. Defend the cutoff, and the objection loses most of its force. Defend the average, and you will keep hearing the objection, because it will keep being right.

Where 5Profiler stands

Design the payoff out of the instrument and there is less left to argue about. 5Profiler measures personality in formats built to make strategic responding a poor investment, and no scale is read alone: personality enters a composite alongside ability and other role-relevant measures, scored and weighted by rule for every candidate, so no single self-report dial decides an outcome. What a hiring team receives is a ranking that is expensive to climb by pushing one trait, together with a record of how the composite was assembled, which is what a candidate's appeal or a regulator's question actually tests. Role-referenced scoring sits underneath it. No format removes the incentive to look employable; the aim is to make acting on it a poor use of a candidate's effort.

Read the science behind the platform · See it on your roles

References

  1. Birkeland, S. A., Manson, T. M., Kisamore, J. L., Brannick, M. T., & Smith, M. A. (2006). A meta-analytic investigation of job applicant faking on personality measures. International Journal of Selection and Assessment, 14(4), 317–335.
  2. Christiansen, N. D., Burns, G. N., & Montgomery, G. E. (2005). Reconsidering forced-choice item formats for applicant personality assessment. Human Performance, 18(3), 267–307.
  3. Hough, L. M., Eaton, N. K., Dunnette, M. D., Kamp, J. D., & McCloy, R. A. (1990). Criterion-related validities of personality constructs and the effect of response distortion on those validities. Journal of Applied Psychology, 75(5), 581–595.
  4. Morgeson, F. P., Campion, M. A., Dipboye, R. L., Hollenbeck, J. R., Murphy, K., & Schmitt, N. (2007). Reconsidering the use of personality tests in personnel selection contexts. Personnel Psychology, 60(3), 683–729.
  5. Ones, D. S., Viswesvaran, C., & Reiss, A. D. (1996). Role of social desirability in personality testing for personnel selection: The red herring. Journal of Applied Psychology, 81(6), 660–679.
  6. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
  7. Viswesvaran, C., & Ones, D. S. (1999). Meta-analyses of fakability estimates: Implications for personality measurement. Educational and Psychological Measurement, 59(2), 197–210.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.