When a first-line management job opens, the selection instrument is usually a spreadsheet sorted by individual output. The top row gets the interview and generally the offer; rewarding the best salesperson with the team feels both fair and safe. Across 38,843 salespeople at 131 U.S. firms, the doubling of sales that raised a worker's promotion odds by 32% relative to the base rate also predicted a 6.1% decline in each subordinate's sales once that worker became the manager. Selecting managers by sorting the output column converts the best seller into a worse manager, systematically and at scale.
The spreadsheet persists because it answers a real question, just not this one. Individual output is visible, auditable, and hard to argue with, which makes it the default evidence in promotion decisions. But a first-line manager's job is organizing other people's production: setting priorities, coaching, allocating work, absorbing friction. A record of personal sales says almost nothing about any of that, and the single-mindedness that builds such a record may lean against it.
The mismatch would matter less if the decision were rare or reversible. It is neither. First-line managers are the layer through which every frontline employee experiences the company, and a promotion, once announced, is among the hardest decisions to unwind without losing both the new manager and the team's confidence. Organizations that would never make an external management hire on a single number routinely make the internal one exactly that way.
Workplaces have joked about the pattern for half a century: people rise to their level of incompetence. The joke has a name, the Peter Principle, and until recently it lacked a large-scale test. Now it has one, and the test does more than confirm the punch line. It measures the size of the error, exposes the incentive that keeps the practice alive, and points to the selection design that would end it.
Selecting managers by the output column has now been tested at scale
Benson, Li, and Shue (2019) linked sales and personnel records for 38,843 sales workers at 131 U.S. firms from 2005 to 2011, covering 1,553 promotions into sales management, and published the analysis in the Quarterly Journal of Economics under the title the folklore had already written: Promotions and the Peter Principle. In these firms, sales performance dominated promotion decisions. Doubling a worker's sales raised the probability of promotion by 32% relative to the base rate, and being a team's top-ranked seller roughly tripled that base rate.
Then the study follows the promoted sellers into the new job and watches their teams. Each doubling of a new manager's own pre-promotion sales predicted a 6.1% decline in the sales of each subordinate reporting to them. The credential that most improves a seller's odds of promotion points the wrong way once the seller is managing.
Nothing about the result requires the promoted sellers to have gotten worse at anything. The simplest reading is that different jobs reward different profiles, and the ranking that identifies the best performer of the first is uninformative, or worse, about the second. A brilliant closer who hoards accounts, wins by outworking teammates, and treats process as friction is displaying exactly the behaviors a sales contest rewards, and several of the behaviors a management job punishes.
The study also computes what the sorting rule forgoes. Had these firms promoted the workers with the highest predicted managerial value-added, a forecast of how much each candidate would improve subordinate performance, average managerial quality would have been roughly 30% higher. The annotation under Figure 1 is worth sitting with: the gain requires no new hiring and no new pipeline, only the same workers re-ranked by the outcome the role exists to produce.
One further result, which the authors label suggestive rather than conclusive, deepens the diagnosis. Workers whose records showed collaboration experience were promoted less often, yet performed better when they did manage. Held at that evidentiary weight, the finding still stings: the sorting rule may not merely overlook managerial ability but tilt against one of its plausible early markers.
Why does Peter Principle hiring survive its own track record? Because promotion is doing two jobs at once: filling a management vacancy and settling the tournament that keeps salespeople selling. The prize has to be worth chasing, and the gap it motivates is enormous; in these data, a 75th-percentile seller generates roughly 11.5 times the revenue of a 25th-percentile seller. A firm that visibly promotes its top sellers is protecting that gap.
Firms are not oblivious to the tension. Benson, Li, and Shue find they already put less weight on sales in promotion decisions where teams are larger, where a weak manager's reach multiplies, and where frontline incentive pay is stronger, where cash already carries the motivation. The trade-off is real. It argues for separating the vacancy from the tournament rather than for continuing to sort by the wrong column.
A better boss adds more than another worker does
How hard organizations should fight for good promotion decisions depends on how much managers matter, and for years that question ran on anecdote. Lazear, Shaw, and Stanton (2015) answered it with field data, published in the Journal of Labor Economics under a title as plain as its finding: The value of bosses. By matching workers' measured output to the supervisors they worked under, the study isolates what a boss adds to the people below them. The design is what gives the estimates their force: the same workers, observed under different supervisors, produce measurably different output, so the difference belongs to the boss.
The estimates that follow set the scale. Replacing a boss in the bottom 10% of quality with one in the top 10% raises a nine-member team's output by more than hiring an additional worker would. And the average boss contributes about 1.75 times as much to output as the average worker. In Figure 2, the intuitive lever for a struggling team is headcount, and it is the smaller one.
A further finding compounds those: better bosses reduce worker exit. A well-chosen first-line manager improves the team twice, through the output of the people currently on it and through how many of them are still on it next year. Choose badly and the losses stack the same way, in output first and then in the departures that follow it. Exit is also where the damage hides longest. Output declines show up in the next quarter's numbers; departures read as individual choices, spread across months, each with its own story, rarely traced back to the promotion that preceded them. A team that loses the manager lottery seldom files a complaint. It files resignations, one at a time.
Practitioners have arrived at the same alarm by their own route. Gallup's estimate, offered as the firm's own practitioner estimate and not as a peer-reviewed figure, is that companies "fail to choose the candidate with the right talent for the job 82% of the time," and that about one in 10 people possess the talent to manage (Beck & Harter, 2014). Whatever precision one grants those numbers, they describe the machine the field studies measured: a selection process pointed at the wrong evidence, feeding a role whose quality moves whole teams. And the companion estimate is the more useful one for design: if the talent to manage is genuinely uncommon, a process that never measures it will mostly miss it.
Managerial potential is measurable, and it is not charisma
If the output column is the wrong predictor, the natural question is what the right one looks like, and the selection literature has held an answer for decades: measure the person, against the job. Personality is the best-documented starting point. In the classic meta-analysis of traits and leadership, the Big Five jointly reach a multiple correlation of .48 with leadership, the five traits predicting together, with extraversion the strongest single trait at .31 (Judge, Bono, Ilies, & Gerhardt, 2002); this series unpacks the trait framework in its briefing on the Big Five and the extraversion result in its briefing on extraversion at work.
Judge and colleagues report a nuance that argues directly against promoting on presence. In the multivariate model, conscientiousness carried the largest standardized beta, the largest unique contribution once the other traits are held constant. Strip away the overlap among traits, and the disposition that most distinguishes leaders is the tendency to plan, follow through, and impose order, which is a fair sketch of what a first-line manager does all day. A committee dazzled by the confident talker is weighting the single correlate; the trait evidence favors the organizer.
The charisma heuristic persists for the same reason the sales sort does: visibility. Confidence in meetings, like a number in a column, is easy to observe and easy to defend in a debrief, while the behaviors that actually move a team, consistent follow-through, fair allocation of work, early correction of drift, are distributed across hundreds of small moments no committee attends. Personality assessed at facet depth reaches the part of the job a committee cannot watch: a facet score for follow-through stands in for a debrief's recollection of who sounded convincing in the room, and unlike the recollection it can be compared across candidates.
Hogan and Kaiser (2005), reviewing what the field knows about leadership, supply the criterion this evidence should be judged against: leadership is about the performance of teams. Personality predicts it, "who we are is how we lead," and the same measures that forecast it can serve selection and development alike. The criterion matters because it relocates the question. A promotion committee asking "who performed best?" is scoring the candidate; a committee asking "whose team will perform best?" is scoring the manager, and only the second question matches the job being filled. That reframing is also what makes managerial potential assessable in advance, since a criterion stated in terms of team outcomes can be predicted by measures taken before anyone holds the title.
Framed that way, the promotion is revealed as a job change, not a reward for the job being left. Producing and organizing collective effort draw on different behaviors, and a record of the first is silent about the second. First-time manager selection goes wrong at exactly this joint: the evidence that exists in abundance describes a role the candidate is about to stop doing.
Hiring managers from outside does not escape the problem
There is a tempting exit. If internal promotion machinery is this unreliable, skip it: let the outside market produce managers and hire them ready-made. Bidwell (2011) followed that route through the personnel records of a U.S. investment-banking division from 2003 to 2009 and measured what it delivers: external hires into comparable jobs were paid roughly 18–20% more than workers who reached the same jobs by internal moves, received significantly lower performance evaluations across their first two years, and left at higher rates.
DeOrtentiis, Van Iddekinge, Ployhart, and Heetderks (2018) found the same shape of result at the other end of the labor market. In a quick-service retail organization, internally selected managers showed higher job performance than externally hired ones, at lower starting salaries. One study covers investment bankers and the other covers quick-service retail managers, which is some assurance that the pattern is not the artifact of a single industry's peculiarities. Figure 3 plots both comparisons as the published record supports them, with the compensation gap drawn as a range and the evaluation gap drawn as direction only.
External hiring is sometimes the right call, and it still offers no exemption from selection. An external candidate arrives with less verifiable evidence, at a premium. An internal candidate arrives with years of observable behavior that most organizations never convert into anything measurable. The premium itself has a rational core: outsiders carry claims no one can verify, so the market charges for the uncertainty and the first two years are spent finding out. What is harder to defend is the asymmetry inside the same firm: structured processes for strangers, folklore for colleagues, and the richer evidence left uncollected. Whichever way the choice between promoting and recruiting lands, someone still has to assess managerial potential, and the organization holding the richest evidence about a candidate is the one that candidate already works for.
Selecting managers works when promotion is treated as selection
The remedy follows from the diagnosis: run the promotion the way a serious organization runs a hire. That begins with defining the manager role's demands as a written role profile covering what the first-line job in this organization actually requires, from prioritizing a team's work to delivering unwelcome feedback. A profile of the role gives measurement something to aim at; the platform overview shows how such a profile carries a definition into scoring.
Against that profile, internal candidates should face measured predictors (personality at facet depth, judgment, integrity: the evidence base this series documents) and structured internal interviews scored the same way an external candidate's would be; the briefing on running structured interviews covers the mechanics. Interviewing a colleague of eight years feels absurd until the alternative is described plainly: deciding a management job on reputation and a sales ledger. Structure matters more for internal candidates than for outside ones. The interviewers already know the candidates, sometimes for years, and an unstructured conversation collapses into confirmation of existing reputations; scored questions about realistic team situations, delivered identically to every candidate, are the simplest defensible correction available.
The file is completed by evidence the ledger ignores. One kind is collaboration-shaped work, mentoring, onboarding new hires, coordinating across teams, weighted with the caution the researchers themselves apply: the finding that collaborative workers made better managers remains suggestive, not settled. The other is the assessment record the organization already owns. A profile collected at hiring keeps its relevance long after the offer, and the briefing on assessment data after hiring shows how carrying it into development lets managerial potential surface years before a vacancy forces the question.
The element that frees all the others is structural: a parallel track that pays top individual contributors at parity, so that excellence has a reward that does not arrive with direct reports attached. Organizations resist the dual ladder on the theory that status must map to headcount, but the alternative is the tournament the field data describe, one in which the only exit from selling is a job that selling ability does not predict. The pressure Benson, Li, and Shue document relaxes once the prize for selling is money and standing, and management becomes a different job with its own door. And because a new first-line manager resets what a team's day-to-day culture rewards, the reasoning in the briefing on culture fit versus culture add applies with full force here: the person selected will not just inhabit the team's culture but author it.
The audit that tells you whether you have this problem
Start with an audit of the last two years of promotion decisions. For each new manager, ask what evidence beyond individual output the decision used, then pull the team's subsequent performance and retention. This is Figure 1 rerun on your own data, and it establishes whether the problem is hypothetical before anyone is asked to change anything. The audit needs no new tooling; the promotion dates, each new manager's prior output rank, and four quarters of team numbers on either side of the change are almost always sitting in the HR system already. Either outcome is worth having. If the teams under output-selected managers held their performance, the local context differs from the studied one and the burden shifts to explaining how; if they slipped, the case for redesign is already written in your own numbers.
Next, write the manager role profile now, while no vacancy exists and no candidacy distorts the drafting. Selecting managers goes wrong earliest at the definition stage: a committee that has never specified what the job demands will default to the evidence at hand, and the evidence at hand is the output column.
Then commit to process parity. Internal candidates go through the same instruments, the same structured interviews, and the same scoring that an external search would impose. Familiarity with a candidate is genuine information, but ungoverned familiarity is politics with better anecdotes; parity is what makes the internal evidence advantage real rather than rhetorical. It also protects the internal candidates themselves: a promotion granted through the same gauntlet an external hire would face needs no defending afterward, to the team or to the candidates passed over. For organizations moving people across many units at once, that consistency is precisely what an enterprise deployment exists to hold steady.
Finally, measure the process itself. New managers' team performance and retention in the first year are the criterion for judging first-time manager selection; offer-acceptance and time-to-fill describe how fast the process runs and say nothing about how well it chooses. A promotion process that never checks its own outcomes has no way to notice it is failing, which is how the practice this article documents survived long enough to be studied.
The promotion spreadsheet will survive too, and it should. It is an accurate record of who sold the most, and bonuses and recognition should keep flowing along it. The one task it cannot perform is seeing the job on the other side of the promotion. Your best salesperson deserves a close reading, but on the right question, because the column that proves someone can sell was never evidence that they can manage.
Where 5Profiler stands
When a management vacancy opens, 5Profiler reads internal candidates against a role-referenced profile of the manager job itself, so the decisive evidence describes the work ahead, not the work being left behind. Because external candidates are scored against the same profile, an internal shortlist and an outside search finally compete on identical measures instead of on reputation against résumé. And because assessment profiles persist beyond the hiring decision, managerial potential can surface early: a contributor assessed on the way in can be re-read as a management candidate the day a team needs one, with the measurement already in hand.
References
- Beck, R., & Harter, J. (2014). Why great managers are so rare [Practitioner article]. Gallup Business Journal.
- Benson, A., Li, D., & Shue, K. (2019). Promotions and the Peter Principle. Quarterly Journal of Economics, 134(4), 2085–2134.
- Bidwell, M. (2011). Paying more to get less: The effects of external hiring versus internal mobility. Administrative Science Quarterly, 56(3), 369–407.
- DeOrtentiis, P. S., Van Iddekinge, C. H., Ployhart, R. E., & Heetderks, T. D. (2018). Build or buy? The individual and unit-level performance of internally versus externally selected managers over time. Journal of Applied Psychology, 103(8), 916–928.
- Hogan, R., & Kaiser, R. B. (2005). What we know about leadership. Review of General Psychology, 9(2), 169–180.
- Judge, T. A., Bono, J. E., Ilies, R., & Gerhardt, M. W. (2002). Personality and leadership: A qualitative and quantitative review. Journal of Applied Psychology, 87(4), 765–780.
- Lazear, E. P., Shaw, K. L., & Stanton, C. T. (2015). The value of bosses. Journal of Labor Economics, 33(4), 823–861.