Research Personality science

The dark side of strengths.

Most trait–performance relationships are curves, not lines, but hiring treats them as lines and rewards the extremes. Where more stops being better, and how the strength that got someone promoted becomes the flaw that ends it.

A promotion case and an exit review, written about the same person at two points in one tenure, often describe the same behavior twice. The promotion case calls it decisive. The exit review calls it reckless. Thorough becomes obstructive, confident becomes uncoachable, driven becomes the reason three good people resigned. Nobody lied in either document, and the person did not change between them. What changed was how much of the trait the job could absorb, and that translation is where most leadership derailment actually happens.

Selection and promotion systems are built as though this never occurs. Scores are ranked, cut-offs are set at the bottom, shortlists are drawn from the top, and the implicit model behind every one of those operations is a straight line: more conscientiousness is better than less, more assertiveness is better than less, more confidence is better than less. The model is rarely stated, because it was never argued for. It is a byproduct of how scores get used.

The research says otherwise, and it says so from three separate literatures with different criteria and different samples. Conscientiousness levels off and then declines at the top of the scale. Assertiveness impairs leadership at both ends. Narcissism relates to rated leadership effectiveness through an inverted U. The shape recurs often enough to state as a working rule for anyone who buys or runs assessment: the trait that earns the promotion, applied harder under pressure and past the range the role can use, is frequently the trait that ends the tenure.

The leadership derailment base rate that no linear model explains

Start with the number that should have prompted the question decades ago. Reviewing the leadership derailment literature, Hogan and Hogan (2001) report that estimates of the base rate of managerial incompetence run from roughly one-third to two-thirds of managers. That is an enormous span, and the temptation is to quote whichever end suits the argument. The width deserves better treatment than that, because the width is itself informative.

Estimates disagree because the underlying question is defined differently in each study. Some count managers who are fired, demoted, or pushed sideways. Others count managers rated as failing by their own subordinates, which catches people whose organizations have not yet acted. Samples differ too, drawn from different industries, seniority levels, and eras. Studies define failure differently, which is what produces a span that wide. Any single figure taken out of that spread and printed as the derailment rate is a piece of false precision.

Held as a range, the estimate still does real work. It says that managerial failure is a common event rather than an unlucky exception, and it says so across every definition anyone has applied. A finding that survives that much variation in method is usually pointing at something structural. Even the low end of the band is difficult to reconcile with the way managers are chosen, because organizations do not promote at random. They promote people who have performed, who interview well, and who increasingly clear a screening battery. For an organization deciding how much to invest in assessment at scale, the operative question is not which number is correct but whether the current process contains any mechanism that would surface this failure mode before the promotion rather than in the exit review.

One explanation is that assessment simply has weak validity, which would make managerial failure a measurement problem. That reading sits badly with the observation that these failures concentrate among the very people the process actively preferred. A better explanation is that a high score carries more than one meaning, and that the region of the scale where the meaning changes has never been plotted.

Nobody chose the straight line

The linear assumption is worth stating explicitly, because it hides inside operations that look purely administrative. A percentile rank presumes that a higher percentile is a better candidate. A cut-off presumes that risk lives below a threshold and nowhere above it. A weighted composite presumes that each additional point of a trait adds the same increment of expected performance as the point before it. Sorting a shortlist descending presumes the whole scale points one way.

Not one of those operations was adopted after someone examined the shape of the trait-outcome relationship. They are inherited conventions from an era when scores were read off a sheet and the practical question was who to interview next. The convention then hardened into software, and software makes assumptions invisible: an applicant tracking system that sorts by score is making a claim about functional form every time it renders a list, and no one on the hiring panel experiences it as a claim.

There is a statistical version of the same blindness. If a relationship is genuinely curved but is fitted with a straight line, the line will usually come out positive and modestly sized, because most of the sample sits in the range where the curve still rises. The fitted slope averages a strong positive relationship in the middle of the distribution with a flat or negative one at the top. Nothing in the output announces that the top region behaved differently. You have to go looking for the curve, and until recently few analyses did.

Most of these relationships are curves

The looking has now been done, in several places at once. Le and colleagues (2011), publishing in the Journal of Applied Psychology under the title "Too much of a good thing," tested curvilinear relationships between personality traits and job performance directly rather than assuming linearity. Conscientiousness and emotional stability both showed inverted-U relationships: performance rose with the trait, leveled off, and then declined at high trait levels. The pattern was more pronounced in low-complexity jobs than in high-complexity ones, which bounds how far the finding travels.

Ames and Flynn (2007) looked at assertiveness and found the same geometry on a different criterion. Both too little assertiveness and too much were associated with lower leadership effectiveness, and at each extreme the corresponding excess or deficit was the weakness raters cited most often about that leader. Too little was the leading complaint at the low end, too much was the leading complaint at the high end, and neither of those problems is visible in a score that only asks how much assertiveness a person has.

Grijalva and colleagues (2015) carried out a meta-analytic review of narcissism and leadership and reported a curvilinear, inverted-U relationship with leadership effectiveness. Moderate levels related to higher effectiveness than either very low or very high levels. It is the least intuitive of the three results and the most instructive, because a trait with an unambiguously negative reputation still turns out to have a productive middle rather than a monotone downward slope.

Three findings do not make a law, and they should not be stacked as though they replicated one another. They used different criteria, different measurement traditions, and different analytic approaches, and they converge on a shape rather than on a value. What makes the convergence interesting is that none of the three was designed to confirm the others. Each research team tested a specific trait against a specific outcome and found the same departure from linearity, which is the pattern you would expect if the departure is a property of how traits operate in work rather than an artifact of one method.

Read Figure 2 for its one structural claim rather than for its numbers. Three research programs, working on different traits with different outcome measures, all found that the relationship bends. A ranking procedure that sorts candidates from high to low draws its shortlist from the far right of every one of these curves, the region where all three sources report the relationship falling away, whatever the peak's true position. A hiring process can be doing exactly what it was designed to do and still be recruiting from the stretch of the range where the evidence has turned against it.

Why the same trait changes its name

Curves invite a mechanistic question about what actually happens to a person as a trait passes from useful to costly. Judge, Piccolo, and Kosalka (2009) supply the conceptual frame in their review of the leader trait paradigm. The traits that predict who emerges as a leader are not the same set that predicts who is effective once there, and a given trait can be adaptive or maladaptive depending on its level and on context. Assertiveness gets you the room's attention, which is an emergence outcome. Whether it produces a functioning team is a separate question with a separate answer.

That gap explains a pattern most executives recognize. Selection systems, interviews included, are efficient at detecting emergence-related traits, because emergence is what a selection process gets to observe. The room notices the confident, articulate, forceful candidate. Effectiveness is observed later by different people under different conditions, and by then the hiring decision has been made and defended.

The behavioral half of the mechanism is more concrete. Under pressure, people do more of what has worked for them. The manager whose thoroughness produced early wins responds to a crisis with additional thoroughness; the manager whose decisiveness built a reputation responds by deciding faster and consulting less. A strength is a well-practiced habit, and habits intensify under load. That is precisely when the trait is least likely to fit the situation, because the situation has changed and the response has not.

Kaiser and Overfield (2010) put a measurement point on this. Working with ratings of leader behavior, they observe that "too much" of a behavior is reported about as often as "too little," and that conventional rating scales running from low to high have no way to express overuse. A leader who overdoes something registers, on an ordinary instrument, as a leader who has a lot of it, which the instrument's own scoring rule then reads as good news. The failure is not in the rater's perception. It is in a response format that has only one direction of concern.

A domain score cannot show you an overused strength

Look again at the middle column of Figure 3. Every entry is a combination of a high facet and a low one, and that structure is not decoration. It is the reason this whole class of risk is invisible in most reports.

Conscientiousness as a domain is an average over narrow traits that behave differently from one another. Dudley and colleagues (2006), in their meta-analytic work on conscientiousness and job performance, showed that the narrow components carry different relationships with different outcomes rather than acting as interchangeable indicators of one thing. Averaging them produces a number that cannot express "very high achievement-striving alongside notably low deliberation," which is precisely the profile that shows up in the exit reviews. Two candidates can post the same domain score with opposite risk profiles, an argument the companion article on facet-level measurement works through in detail and does not need repeating here.

The consequence for leadership derailment risk is direct. A high domain score is a genuinely ambiguous signal: it might mean broad, moderate elevation across six facets, which is usually the good case, or it might mean two facets near the ceiling with the balancing ones near the floor, which is the case that reads as reckless in eighteen months. A report that stops at the domain forces the panel to guess which one they are looking at, and a panel with no basis to distinguish them will usually read a high score as the strength it was labeled as.

Facet resolution also fixes the interpretation of the curve results. When Le and colleagues found that performance declines at high conscientiousness, the plausible reading is not that carefulness becomes harmful in itself but that the extreme region of the domain is disproportionately populated by unbalanced facet profiles. The article on conscientiousness at work covers the trait's overall validity picture. What matters here is that an average cannot express the combination, and you cannot manage a risk you have averaged away.

Score the range, not the maximum

The practical response is a change in what a report is for. A report built to rank tells you who scored highest. A report built to inform tells you where each candidate sits relative to the range a specific role can use, and flags distance from that range in either direction.

Before describing what that looks like, one boundary needs stating plainly, because this territory is easy to misread. Trait extremity is a role-fit question and not a clinical one: nothing in this literature diagnoses anything, none of these traits is a disorder, and a personality instrument used for hiring has no business implying otherwise. The claim is bounded and modest. A score far outside the range a role can absorb carries risk specific to that role, and the report should say so rather than presenting the same score as a triumph.

Four changes follow from Figure 4, and they are changes to reporting and process rather than to the science. First, define the usable range for each role before candidates are scored, so that the band is a stated expectation and not a post-hoc rationalization of a preferred candidate. Second, flag distance from the band in both directions, which means an instrument whose output can express overuse at all. Third, treat a flag as a question rather than a verdict; the appropriate response to Candidate B is a targeted probe into how that trait has behaved under pressure, not a rejection. Fourth, pair the trait profile with structured behavioral evidence, since the flag identifies where to look and structured interviewing is what does the looking.

The obvious objection is that the bands have to come from somewhere, and that a badly chosen band is worse than no band at all. That is fair, and it sets a standard rather than blocking the approach. A usable range should be derived from the role's actual demands and its known failure modes, documented before scoring, and reviewed when the role changes. Where an organization has incumbent data, the range can be checked against how current performers are distributed. Where it does not, the band is a stated hypothesis that hiring outcomes will confirm or move, which is still a considerable improvement on an unstated hypothesis that more is always better.

The third point is where most implementations go wrong. Extremity is a risk indicator with a wide error band around any individual case, and plenty of very high scorers perform well, particularly where the role genuinely rewards the trait without limit. Reading a flag as a disqualification substitutes one mechanical rule for another. The article on interests versus ability makes a parallel argument about signals that inform a decision without deciding it.

The promotion decision is where this costs the most

Everything above applies to hiring, but the base rate lives in promotion, and promotion is where organizations have the least discipline. External candidates get assessed. Internal candidates get discussed, and the discussion runs on the promotion-case vocabulary in the left column of Figure 3: a record of results plus a set of adjectives that describe the very trait extremity that will become the problem.

Three specific corrections are worth the effort. Assess internal candidates on the same instrument as external ones, so that a promotion decision has trait data at all rather than only performance history. Define the target role's usable ranges separately from the current role's, because the ranges move; the deliberation that a senior individual contributor can safely run low on is often exactly what the manager role requires. And treat a strong track record as evidence about the role the person has been doing, which it is, rather than as evidence about the role they are being moved into.

The economics favor this more than the effort suggests. The cost of a leadership derailment plausibly scales with seniority, since failure at that level takes longer to surface, damages more relationships, and is harder to unwind, which is an argument for spending the assessment effort there. The expensive mistakes in most organizations are not the low scorers who slipped through. They are high scorers promoted on the strength of a trait nobody bounded.

The corrections above wait on no new science. The curvilinear findings are more than a decade old, the leader trait framework older still, and the measurement fix, reporting at facet level against role-specific usable ranges, is available today. What has to change is the assumption embedded in every descending sort: that the candidate at the top of the list is the candidate the role can best use. Sometimes that is true. The evidence says it is not reliably true, and a process that has never asked which case it is facing has no way to tell the difference.

Where 5Profiler stands

Read a 5Profiler report from the middle outward rather than from the top down. Each facet is read against the role's usable range, not a population norm, and 5Profiler flags distance from that range in both directions, so a score far above the band arrives as a question to put to the candidate rather than as a reason to move them up the list. Because the report resolves personality at facet level, the combinations this article is about, high Achievement-Striving alongside low Deliberation among them, stay visible instead of averaging into one reassuring domain number. You can see that format on a sample report.

Read the science behind the platform · See it on your roles

References

  1. Ames, D. R., & Flynn, F. J. (2007). What breaks a leader: The curvilinear relation between assertiveness and leadership. Journal of Personality and Social Psychology, 92(2), 307–324.
  2. Dudley, N. M., Orvis, K. A., Lebiecki, J. E., & Cortina, J. M. (2006). A meta-analytic investigation of conscientiousness in the prediction of job performance: Examining the intercorrelations and the incremental validity of narrow traits. Journal of Applied Psychology, 91(1), 40–57.
  3. Grijalva, E., Harms, P. D., Newman, D. A., Gaddis, B. H., & Fraley, R. C. (2015). Narcissism and leadership: A meta-analytic review of linear and nonlinear relationships. Personnel Psychology, 68(1), 1–47.
  4. Hogan, R., & Hogan, J. (2001). Assessing leadership: A view from the dark side. International Journal of Selection and Assessment, 9(1–2), 40–51.
  5. Judge, T. A., Piccolo, R. F., & Kosalka, T. (2009). The bright and dark sides of leader traits: A review and theoretical extension of the leader trait paradigm. The Leadership Quarterly, 20(6), 855–875.
  6. Kaiser, R. B., & Overfield, D. V. (2010). Assessing flexible leadership as a mastery of opposites. Consulting Psychology Journal: Practice and Research, 62(2), 105–118.
  7. Le, H., Oh, I.-S., Robbins, S. B., Ilies, R., Holland, E., & Westrick, P. (2011). Too much of a good thing: Curvilinear relationships between personality traits and job performance. Journal of Applied Psychology, 96(1), 113–133.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.