Research Personality science

Agreeableness cuts both ways.

In team-based work agreeableness is the strongest personality predictor (.33); across four studies it also predicts lower earnings, and in zero-sum bargaining it is a measured liability. Context decides the sign.

The same employee can be the most valuable person in one room and a liability in the next. In the morning’s project retrospective, the team’s most agreeable engineer keeps the meeting from curdling: criticism gets restated as fixable, the defensive colleague stays in the conversation, and the group leaves with its working relationships intact. In the afternoon’s supplier negotiation, the same engineer treats the other side’s opening number as a reasonable place to start, and concedes margin the company will miss. Nothing changed on the walk between rooms except the task, and the value of agreeableness at work changed with it.

Both performances come from one trait, working as designed, and no other Big Five dimension carries the contrast so far. The average validity of agreeableness at work across all jobs is a forgettable .07, a validity coefficient (the correlation between assessment scores and later job performance) that most selection texts mention once and move past; the flagship briefing on what predicts job performance explains how such estimates are built. In jobs organized around sustained teamwork, that same disposition is the strongest personality predictor in the set, at .33. Across close to 4,000 workers in three national samples, it predicts lower income. Across a distributive bargaining table, it is a measured liability. Same trait, opposite signs, all of it in the published record.

So the productive question for a hiring team is not whether agreeableness is good. It is what the role’s interactions are structured to reward. Sustained interdependent work returns the trait’s value in performance; zero-sum advocacy penalizes it; brief, scripted service barely registers it. The rest of the evidence sorts itself by that principle, and it ends in a scoring practice most hiring processes skip.

The trait is a motive, and the task decides its sign

In the research literature, agreeableness is defined by what it wants, not by what it can do. The standard handbook treatment characterizes it as the dimension of personality organized around the motivation to maintain positive relations with others (Graziano & Eisenberg, 1997). Everything paradoxical about the trait follows from that definition. A motive travels with the person into every situation; whether it helps depends on whether the situation rewards what the motive keeps trying to do.

In a project retrospective, maintaining relationships is a large share of the work itself, and for long stretches the whole of it. Friction is the characteristic failure mode of interdependent work, and a person dispositionally driven to absorb friction, to extend the benefit of the doubt, to repair small ruptures before they widen, contributes performance every time the motive engages. In a distributive negotiation that motive competes with the task, because claiming value strains a relationship by design. A negotiator built to protect the relationship gives ground exactly where the task requires ground to be held.

The distinction between motive and skill also matters for how the trait gets judged in hiring. Interview panels routinely read warmth as collaboration ability, but a motive is a direction of effort, and direction is hard to observe in a conversation with no real task attached. A candidate low on agreeableness can execute cooperation deliberately when the stakes are visible, and a highly agreeable one can cooperate a project into the ground, accepting scope the team cannot carry and withholding the objection that would have saved a quarter. What the trait score forecasts is where effort defaults when nobody is deciding consciously, which is valuable information about long-run behavior and nearly invisible in an interview.

Scripted one-to-one service sits between the two. Exchanges are brief, courtesy is standardized, and the relationship resets with each customer, so dispositional differences in relationship maintenance have less variance to explain. One motive, three task structures, three different returns. The evidence attaches numbers to each in turn.

The all-jobs validity of agreeableness at work is an average of opposites

The number that made the trait forgettable comes from Barrick and Mount (1991), the Personnel Psychology meta-analysis that organized the modern study of personality and job performance. Pooled across occupational groups, the validity of agreeableness came to .07, and pooling across criterion types produced an identical figure. The criterion detail is scarcely livelier: .06 against job proficiency, .10 against training performance, .14 against personnel data. Conscientiousness posted .22 and generalized across every occupational group studied, and that contrast set the hierarchy hiring practice absorbed: one personality trait worth measuring everywhere, four traits worth measuring nowhere in particular. The conscientiousness briefing in this series traces what happened to the stronger number; this article is about what the weaker one conceals.

An average is only as meaningful as the homogeneity of what it averages, and for the agreeableness–job performance relationship the homogeneity assumption fails. The all-jobs figure pools settings where the trait predicts strongly with settings where it barely registers, and the wider literature adds settings where its outcomes run negative outright. A coefficient near zero is therefore ambiguous between two very different worlds: a trait that never matters, and a trait that matters in opposite directions depending on where it is standing.

The general case is a predictor whose validity is conditional on a feature of the situation: pooled over situations, it will always understate itself, and agreeableness is the Big Five’s starkest example. The practical consequence runs in both directions at once. A reader trusting the all-jobs table will pass on a trait their team roles need, and a reader generalizing from team evidence will overweight it in roles built for advocacy. Neither error announces itself, because both are backed by a real coefficient from a real meta-analysis.

Measurement choices compound the averaging problem. The 2022 re-estimates put generic agreeableness at .10 against job performance and work-contextualized measures at .19 (Sackett et al., 2022), a measurement gap the extraversion briefing in this series works through in detail. For agreeableness at work, more than for most traits, validity moves with how the question is framed, and a measure that never mentions the job hides the very contrasts the rest of this evidence maps.

In team jobs, agreeableness outpredicts every other Big Five trait

The decisive evidence comes from Mount, Barrick, and Stewart (1998), a meta-analysis in Human Performance built entirely from jobs involving interpersonal interactions. Its contribution was to refuse to treat those jobs as one category. Studies were sorted by the structure of the interaction: team jobs, where work runs through sustained, interdependent relationships among a fixed set of colleagues, and dyadic service jobs, where work is a stream of short, largely transactional exchanges with customers or clients.

The sort separates what the all-jobs average had blended. In team jobs, agreeableness reached a corrected validity of .33 across four studies covering 678 workers, which made it the strongest personality predictor in that setting; emotional stability tracked that gradient at .27, a parallel the emotional stability briefing in this series picks up. In dyadic service jobs, agreeableness managed .13 across seven studies and 908 workers, and emotional stability slipped to .12. Figure 1 plots the spread, with the pooled interpersonal estimate and the all-jobs average marking the reference points between and below.

The criterion side of that analysis ties the result to the mechanism. When the outcome being rated was specifically how well an employee handled interactions with others, the validity of agreeableness rose to .35, its best showing anywhere in the study. The trait predicts best where the outcome sits closest to what the underlying motive is for; the further the criterion drifts from relationship maintenance, the less the motive resembles ability.

A caveat keeps the number in proportion: the team-jobs estimate rests on four studies and 678 workers, a modest base for a finding this consequential. What makes it credible is not its size but its coherence. Emotional stability moves in parallel across both categories, the criterion analysis peaks where the motive account says it should, and group-level research arrives at that conclusion from a different direction. The category has also outgrown its 1998 examples: cross-functional product squads, operating-room teams, shift crews handing work across a boundary twice a day. Any role where a fixed group must coordinate repeatedly under friction sits in the cell the .33 came from, whatever the industry label says.

That group-level convergence comes from composition research. In field teams, performance tracks the group’s agreeableness composition, down to the influence of its least agreeable member (Bell, 2007); the team personality composition briefing in this series works through those numbers and what they imply for assembling groups rather than screening individuals one at a time. For teamwork-heavy hiring, the personality evidence rarely comes this well sorted.

Zero-sum bargaining converts the motive into a liability

Barry and Friedman (1998) supplied the cleanest demonstration that the sign can flip. Writing in the Journal of Personality and Social Psychology, they examined how bargainer characteristics played out across the two canonical negotiation structures: distributive bargaining, a fixed-sum contest over a single issue such as price, where whatever one side gains the other loses, and integrative negotiation, a multi-issue problem where trades between issues can enlarge what both sides take home.

In the distributive structure, agreeableness behaved as a measurable liability. Agreeable negotiators were more susceptible to the pull of the opening offer, an anchoring effect of −.10 in standardized terms, and more agreeable sellers captured less of the available gain, at −.12. Extraversion carried a parallel penalty in distributive bargaining, which locates the liability in the structure of the task and not in one trait.

In the integrative structure, the penalty vanished. No personality trait predicted worse joint outcomes, and the variable that governed the size of joint gains was cognitive ability, whose standing as a predictor the general mental ability briefing in this series examines. The reversal is theoretically tidy. An integrative problem is solved by understanding the counterpart’s interests, sharing information, and finding trades, activities a relationship-maintenance motive supports. A distributive contest is won by withholding information, holding anchors, and tolerating a counterpart’s displeasure, which reads like a list of what agreeable people are motivated to avoid.

Mapping the laboratory contrast onto real roles takes some care, because few jobs are purely one structure. Procurement, claims settlement, and rate renegotiation live mostly on the distributive side. Account management, partnerships, and customer success are integrative most of the year. A sales cycle switches structures midstream: integrative while the solution is being shaped, distributive in the closing session where terms get fixed. A role audit that asks how often the job requires claiming value against a counterpart, and how much of the year’s outcome those episodes decide, gives the trait weighting an empirical footing that a job title never will.

The study’s most useful finding for practice is the buffer. The distributive penalty concentrated among negotiators who entered without high aspirations; those who set ambitious targets before the session showed no measurable trait liability. A target fixed in advance does the resisting that the disposition declines to do. Figure 2 traces both structures, with the aspiration buffer marked where it operates.

The trait is costliest at the moments when rewards get claimed

Negotiations are episodes; income is a career-long accumulation, and the trait’s fingerprints show up there too. Judge, Livingston, and Hurst (2012) examined agreeableness and income across three national samples, 560 workers in the NLSY97, 1,681 in MIDUS, and 1,691 in the WLS, and found one result repeated in each: higher agreeableness predicted lower earnings, with standardized effects between −.12 and −.14, and the estimates held with extensive statistical controls in place. The question in the paper’s title, whether nice people at work really do finish last, gets an uncomfortable answer wherever last is measured in income.

The fourth study located part of the mechanism. In an experiment with 460 participants, evaluators reviewing candidate profiles were less likely to recommend the agreeable candidates for a management track. So the earnings gap does not originate only at bargaining tables. Part of it forms earlier, in judgments about who belongs on the track that leads toward management.

The design of the studies bounds the interpretation. The three surveys record income, not the quality of anyone’s work; the experiment records an observer’s recommendation, made early and on limited exposure, in exactly the kind of setting where self-advocacy reads as leadership potential. Together they describe a career-long filter that keeps meeting the trait in its worst settings, the moments when rewards are claimed, and keeps missing it in its best ones, the years of interdependent work in between.

Leadership research completes the pattern. Judge, Bono, Ilies, and Gerhardt (2002), in the standard quantitative review of personality and leadership, put the trait’s overall leadership correlation at .08, the weakest nonzero relationship among the Big Five. Split the criterion, though, and the small average once again comes apart: agreeableness correlates .05 with leadership emergence, whether a person comes to be seen as leader material in a group, and .21 with leadership effectiveness, how well someone actually leads once installed.

That asymmetry maps cleanly onto the two structures this article keeps finding. Emergence is a competition for standing among near-strangers, structurally closer to distributive bargaining than to teamwork. Effectiveness is sustained interdependent work with a fixed cast, the setting where the trait’s validity peaks. Promotion systems that read potential off emergence therefore select against a trait the job itself rewards, a failure the briefing on selecting first-line managers examines in its own right. Figure 3 puts the income and leadership findings on one axis.

Zoom out, and the trait is overwhelmingly an asset

Two adverse findings do not make a defective trait. Wilmot and Ones (2022) aggregated the empirical record on agreeableness for Personality and Social Psychology Review: 142 meta-analyses covering 275 variables and more than 1.9 million participants. Across that record, agreeableness produced effects in the desirable direction for 93% of the variables examined, and teamworking sits among the themes most characteristic of the trait. The base rate is emphatically on the side of hiring agreeable people.

That base rate sets the default a hiring process should start from. Where a role has not been analyzed, the defensible prior is that agreeable candidates will do better across the outcomes an employer cares about, and the burden of proof sits with anyone arguing to screen the trait down. What the exceptions change is not the prior but the conditions under which a hiring team is entitled to depart from it.

The negotiation and income results are exceptions of a specific and manageable kind. They arise where the task structure sets individual claiming against relationship maintenance, a minority of work settings and, more usefully, an identifiable one. Nothing about the trait announces when it has entered such a setting. Finding those settings in advance is design work, and it is work the trait cannot do for itself.

Weight agreeableness by the role’s interaction structure, not by its reputation

Map the role before weighting the trait. The operational question is how much of the job runs through each of three interaction structures: sustained interdependence, where the same people coordinate repeatedly and friction compounds; transactional advocacy, where the employee’s task is to claim value from counterparts trying to do the same; and scripted one-to-one service, where exchanges are brief and courtesy is standardized. The first argues for weighting agreeableness up, on the strength of the team-jobs evidence. The second argues for weighting it down, or for hiring the person anyway and building structure around the bargaining. The third argues for placing the assessment’s weight on other attributes entirely.

Service hiring is where intuition most needs that correction, because “agreeable equals service” feels self-evident and the evidence disagrees. At .13, the trait’s return in dyadic service work is real but modest, far below its team-jobs value, plausibly because the courtesy in such roles is scripted, and a script needs no disposition behind it. Screening service candidates primarily on warmth puts the assessment’s scarcest resource, decision weight, where the evidence least supports it.

Where distributive bargaining is unavoidable, the aspiration buffer converts a personality finding into an operating procedure. Targets set in numbers before the session, with a walk-away fixed in advance, eliminated the measured disadvantage in Barry and Friedman’s data. The logic is familiar from another corner of selection research: structure carries the parts of performance that disposition would otherwise decide, the reason structured interviews outperform unstructured ones (the briefing on running structured interviews covers the mechanics).

The emergence-effectiveness gap reads better as a process problem than as a trait problem. A pipeline that identifies future leaders by who emerges, who speaks first, claims credit comfortably, and self-nominates, selects on the .05 column and then lives with the consequences in the .21 column. Promotion criteria anchored in demonstrated effectiveness inside interdependent work make the trait’s leadership value visible before the appointment instead of after it.

Finally, the measurement itself has to be role-aware. A single domain score for agreeableness blends facets that these structures treat differently, and an all-purpose interpretation of that score repeats, inside one candidate’s report, the averaging mistake the all-jobs validity made across an entire literature. Scoring against a role profile is the corrective, and it is the design argument behind role-referenced assessment platforms.

The colleague from the opening will keep doing both things, holding a team together in one room and giving ground in the other, because traits do not read agendas. Which room the trait works in is the organization’s choice, made at hiring and again at every promotion. The paradox of agreeableness at work dissolves the moment someone checks what the room asks before deciding who to send in.

Where 5Profiler stands

An evidence base where one trait changes sign between rooms needs scoring that knows which room it is in. 5Profiler applies role-referenced scoring at facet resolution, and every report names the role profile a candidate was read against. A team-lead opening and a procurement-negotiator opening are therefore built as two separate profiles, each carrying its own weights, and the same agreeableness result is read against each: weighted up where the interdependence evidence points, flagged for structure where the bargaining evidence does.

Read the science behind the platform · See it on your roles

References

  1. Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26.
  2. Barry, B., & Friedman, R. A. (1998). Bargainer characteristics in distributive and integrative negotiation. Journal of Personality and Social Psychology, 74(2), 345–359.
  3. Bell, S. T. (2007). Deep-level composition variables as predictors of team performance: A meta-analysis. Journal of Applied Psychology, 92(3), 595–615.
  4. Graziano, W. G., & Eisenberg, N. (1997). Agreeableness: A dimension of personality. In R. Hogan, J. Johnson, & S. Briggs (Eds.), Handbook of personality psychology (pp. 795–824). Academic Press.
  5. Judge, T. A., Bono, J. E., Ilies, R., & Gerhardt, M. W. (2002). Personality and leadership: A qualitative and quantitative review. Journal of Applied Psychology, 87(4), 765–780.
  6. Judge, T. A., Livingston, B. A., & Hurst, C. (2012). Do nice guys—and gals—really finish last? The joint effects of sex and agreeableness on income. Journal of Personality and Social Psychology, 102(2), 390–407.
  7. Mount, M. K., Barrick, M. R., & Stewart, G. L. (1998). Five-factor model of personality and performance in jobs involving interpersonal interactions. Human Performance, 11(2–3), 145–165.
  8. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
  9. Wilmot, M. P., & Ones, D. S. (2022). Agreeableness and its consequences: A quantitative review of meta-analytic findings. Personality and Social Psychology Review, 26(3), 242–280.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.