Research Fit, teams & culture

Hire for the roster, not just the role.

Team-average agreeableness and conscientiousness lift performance — and one low scorer can matter more than the mean. The operators, the evidence, and the craft of team-aware hiring.

One hire changes the team's arithmetic: the mean, the spread, and the minimum all move the day someone joins. Selection science is built to measure one person at a time, and the statistics that predict how a team performs belong to the group, appearing on no individual's report. Almost no hiring process computes any of them.

Team personality composition, the study of how a group's trait makeup predicts what the group achieves, has produced a literature with an unusually practical shape. It says which traits matter at the team level, and they are fewer than intuition suggests. It says which statistic of the team carries the effect, and it is not always the average. And it says how much the setting changes the answer, which turns out to be: a great deal.

The blind spot is structural. Selection research validates instruments person by person, against person-level criteria, and within that frame the evidence is strong. But group outcomes depend on coordination, mutual trust, and covering for one another, and those are relational products: they emerge between people, at rates set by who else is in the room. A literature that scores candidates one at a time has no column for that, and the gap is what composition research fills.

The findings ask for no new instruments, only a second reading of scores organizations already collect, taken at the level of the team a candidate is about to join. That second reading changes shortlists, because the same person adds different things to different teams.

A team can be scored four ways, and the work decides which score binds

A five-member team measured on one trait can be summarized several ways. The mean describes level: how much of the trait the group carries overall. The variance describes similarity: whether members resemble one another or sit far apart. The minimum is the lowest member's standing; the maximum is the highest. Team researchers call these aggregation methods, and the choice among them is not a formality. Each one is a different theory of how individual dispositions become collective behavior.

The difference is easiest to feel through the variance. Two teams can share an identical mean on conscientiousness while living in different worlds: one holds five moderately organized people with common standards for deadlines and detail; the other pairs meticulous planners with cheerful improvisers, and its mean describes nobody in the room. Scoring a team by its mean asserts that trait quantities pool. Scoring the variance asserts that alignment among members matters in itself. Scoring the minimum asserts that one person can cap the group. These are competing claims about mechanism, and data can arbitrate them.

The logic for choosing comes from Steiner's (1972) typology of group tasks. In additive tasks, members' outputs pool: total sales across a territory, tickets cleared by a support desk. Contributions sum, so the mean is the natural statistic. In conjunctive tasks, the work passes through every member, as in a relay, a handoff-dense production chain, or a surgical team; the group can be no better than its weakest contributor, so the minimum binds. In disjunctive tasks, one member's solution can carry everyone, the way a single correct proof settles a problem, and the maximum matters most. Figure 1 pictures these statistics on a single roster, with the task logic beneath them.

The typology works as a hypothesis generator. As a lookup table it overpromises. Bell (2007), a meta-analysis of deep-level composition variables in the Journal of Applied Psychology, tested whether the statistics matched Steiner's task types as cleanly as the logic implies and found limited support for the tidy matching. What holds up instead is coarser and more useful: a short list of traits that matter at the team level, and the finding that for at least one of them the weakest member, not the average member, carries the prediction.

Team personality composition is mostly an agreeableness and conscientiousness story

The central estimates come from Peeters, van Tuijl, Rutte, and Reymen (2006), a meta-analysis of personality and team performance published in the European Journal of Personality. It separated questions that casual talk about team chemistry blurs: does the team's level on a trait predict performance (elevation, the mean), and does the team's diversity on that trait predict it (variability, the spread)?

On elevation, two traits carried the signal. Team mean agreeableness correlated .24 with team performance (a corrected correlation, ρ, of the kind used across the selection-validity literature; its 90% confidence interval runs .09 to .39), and team mean conscientiousness .20 (interval .09 to .31). Mean extraversion, emotional stability, and openness added essentially nothing; their estimates sit within rounding distance of zero in Figure 2. The ordering also flips at the team level: individually, conscientiousness is the headline personality trait, but for teams, agreeableness edges ahead.

At the individual level, personality validities the size of .24 and .20 sit mid-table, useful in a composite but rarely decisive alone. At the team level, the unit of outcome is larger: the criterion is what an entire group produces together, so a correlation of this size describes systematic differences in collective output, repeated across every project the team touches. The same coefficient means more when it moves more people.

Variability ran the other way. Spread in agreeableness correlated −.12 with team performance (interval −.16 to −.07), and spread in conscientiousness −.24 (interval −.33 to −.14), the strongest of the variability effects. Read together, the findings draw a specific picture: teams do best when they are collectively high, and mutually similar, on the two traits that govern cooperation and reliability. A team of very conscientious and fairly conscientious members outperforms a team with the same mean built from opposites (Peeters et al., 2006).

Then comes the moderator that matters for anyone hiring: the professional-versus-student split. In samples of professional teams, intact groups doing consequential work, the elevation effects roughly doubled: mean agreeableness predicted team performance at .51 and mean conscientiousness at .42. In student project teams assembled for a semester, the same estimates were roughly .02 and .00. Much of team research is conducted on student teams because students are available; nearly all hiring concerns professional teams. The field numbers are the relevant ones, and the field numbers are large.

The split has a mundane explanation and a consequential one. The mundane: student project teams are short-lived and low-stakes, so there is little shared fate for personality to shape. The consequential: composition effects need time and interdependence to surface. Traits express themselves through repeated interaction, through who volunteers, who follows through, who bristles at feedback, and a team that will exist for years gives those tendencies room to accumulate into performance. Hiring managers staff the long-lived kind.

Figure 2 plots the team personality composition estimates with their confidence bands, and two of its features repay attention. The whiskers on the two variability effects sit entirely below zero: on agreeableness and conscientiousness, spread is a directional drag on performance, not noise around a null. And the lower panel shows what pooling across settings hides: averaged in with semester-long student projects, composition looks like a modest influence; measured where hiring actually happens, it does not.

In real teams, the weakest member can matter more than the average

Bell (2007) asked a different moderator question: do laboratory teams and field teams give the same answer? Deep-level composition, in her vocabulary, is the psychological makeup of a membership, the traits and values members carry, as distinct from the readily visible facts of a roster. And the settings do not agree. In laboratory settings, where strangers assemble for an afternoon, personality composition barely registered. In field settings, where members share history, workload, and consequences, the effects were strong, and the variables at the top of the field list were team minimum agreeableness and team mean conscientiousness, alongside mean openness and team-oriented values such as collectivism and preference for teamwork.

The switch of statistic inside that sentence is the finding. For conscientiousness, field prediction runs through the mean: reliability aggregates, and every dependable member adds something. For agreeableness, it runs through the minimum: the least agreeable member is the number to watch (Bell, 2007). Cooperation, on this evidence, behaves conjunctively even when the task looks additive, because any single member can withhold it.

The mechanism is not mysterious. Cooperative norms hold only while members expect reciprocity, so a single colleague who keeps score, bristles at requests, or treats help as weakness gives everyone else a reason to hedge, and hedging spreads. The team drifts from open handoffs toward guarded, minimal exchange, and that drift is cohesion falling, whether or not anyone measures it.

Barrick, Stewart, Neubert, and Mount (1998) watched that mechanism up close. Across 51 real work teams, 652 employees in all, they scored each team by mean, variance, minimum, and maximum. Teams whose lowest-scoring member was low on agreeableness or emotional stability were less cohesive, and wide variability in conscientiousness went with worse team performance, the same spread penalty Peeters and colleagues later estimated at −.24. Their conclusion was blunt: even one disagreeable member can disrupt a team's cooperation. At the individual level, agreeableness has a far more mixed record, examined in the collaboration-paradox article. At the team level, the record is not mixed at all.

Figure 3 puts one team beside another that matches its average but holds a member near the floor. Averaging launders exactly the information the field evidence says is predictive: report only the mean, and the two teams are indistinguishable; the low scorer vanishes into it. Any dashboard that reports team averages, and most report nothing else, carries the same blindness by design.

Composition shows up in how teams behave before it shows up in results

Prewett, Walvoord, Stilson, Rossi, and Brannick (2009) revisited the team personality literature with attention to what studies used as the criterion, and found that team conscientiousness and agreeableness relate most strongly to team behaviors, how members cooperate, coordinate, and sustain effort day to day, with effects on final outcomes flowing through those processes. They also confirmed that the answers move with the aggregation method and with the team's workflow pattern. For a manager, the translation is diagnostic: composition problems announce themselves in process first, in meetings, handoffs, and conflict, before they reach the quarter's numbers.

Spread is also something staffing can set on purpose. Humphrey, Hollenbeck, Meyer, and Ilgen (2007) call the practice seeding: staffing deliberately to raise or lower a team's variance on a chosen trait. Seeding treats the variability findings as an instrument rather than a warning. If similarity on conscientiousness and agreeableness is an asset, staffing can build it; if a team needs range on some other trait, staffing can build that too, while keeping the added spread off the traits cohesion runs on. Organizations already make this decision implicitly every time they debate whether the next hire into a team of similar temperaments should resemble them or offset them. Seeding makes the decision explicit, trait by trait, with the variability evidence in view.

LePine (2003) adds the twist. In a study of 73 teams, the task changed without warning partway through. Before the change, composition did not separate the teams. After it, teams higher in cognitive ability, achievement orientation, and openness performed better, and teams heavier in dependability performed worse, with the difference running through how quickly members restructured their roles. Conscientiousness splits at this altitude: its achievement side helps a team adapt, while its dependability side, the rule-keeping, routine-preserving side, resists the very reorganization a changed task demands. The individual-level evidence on conscientiousness shows the same facet structure; LePine's teams show what it means for staffing a world that will not hold still. The result warns against over-optimizing for the stable case: a roster tuned tightly to today's process is, on this evidence, the roster likeliest to struggle when the process is redrawn.

Whom to hire for a team is a different question from whom to hire for a role

Profile the destination team before shortlisting into it. Every member completes the same personality measure a candidate would, and the team's mean, spread, and minimum on each work-relevant trait fall out as a one-page profile, the team-level counterpart of the facet profile in a sample individual report. Team composition hiring starts from that document, with the job description as the second input.

Partial coverage still works. A profile built from most of a team's members locates the mean well enough, and it answers the question that changes decisions most often: whether anyone already sits near the floor on the traits this work strains. The statistic with the largest bearing on a hire is also the one that needs the fewest people to compute.

Next comes deciding which statistic the work makes binding, and here Steiner's typology returns as a set of questions to ask of the workflow. Handoff-dense, interdependent work argues for protecting the minimum on agreeableness and emotional stability, because cohesion is gated by the lowest scorer (Barrick et al., 1998; Bell, 2007). Additive production work argues for raising the mean on conscientiousness. Genuinely disjunctive work, where one insight can carry the team, argues for a high maximum. Most real roles mix the three task types, so the practical question is not which category the job falls into but which failure the work forgives least.

Building balanced teams needs the same specificity. As commonly used, the phrase imagines diversity on every dimension at once. The evidence is more particular: on agreeableness and conscientiousness, similarity is the asset, and added spread carries a measured penalty (−.12 and −.24 in Peeters et al., 2006); elsewhere, variance can be seeded in deliberately (Humphrey et al., 2007). Balance is a trait-by-trait decision, not a team-wide aesthetic. When a profile shows a team already wide on conscientiousness, the composition-aware move is to hire toward the team's center on that trait, absorbing spread instead of adding to it; the variability effects are the quantified case for exactly that instinct.

One boundary keeps the practice clean. Team personality composition concerns measured traits, and nothing in this evidence involves, or licenses inferences from, age, gender, or background. The statistics are computed from trait scores, and only from trait scores; a team that looks homogeneous demographically can be wide open on conscientiousness, and the reverse. Conflating the two levels reads difference where the data measure none.

The team is also not the organization. Person-organization values fit asks whether a candidate's values match the company's; composition asks whether this candidate's traits complete this specific group of people. The two screens sit at different levels of the same question, and they can disagree: right company, wrong five people.

The last input is time. For stable, procedure-bound work, a dependability-heavy roster is the reliable build. For work that reorganizes often, LePine's (2003) adaptation results argue for weighting achievement facets over dependability facets when adding members. Whom to hire for a team depends not only on the team's present statistics but on how often its task will be rewritten.

Recompute the team's numbers before the offer goes out

The most direct implication is for shortlists. Two finalists with identical role scores stop being interchangeable once the destination team enters the file: one of them raises the team's minimum on agreeableness, and the other becomes the new minimum. Team-aware shortlisting does not overturn role-based ranking; it re-sorts the top of it. In practice that means carrying the team's statistics alongside each finalist's individual scores and asking, for each candidate, what the team's numbers look like after the join.

A concrete version: a company hires two account managers in the same cycle, for two regional teams, from one finalist pool. One region's profile shows a high, tight cluster on agreeableness; the other shows a solid mean with one member already near the floor. The first team can absorb a middling-agreeableness finalist without moving its minimum. For the second, the same finalist deepens an existing weakness, and a candidate who ranked lower on role scores but higher on the trait may be the stronger hire there. Identical candidates, defensibly different decisions, because the destination statistics differ.

The second implication is procedural: write the operator argument down. If a role sits in a handoff-dense team and the hiring plan therefore protects minimum agreeableness, that reasoning should exist on paper before the search opens, the way structured processes pre-commit to interview questions. Documentation keeps the practice anchored to the workflow, where the evidence lives, and away from impressions about who would be pleasant company, where it does not. It also changes what the organization can say about its own decision afterward, since protecting a team minimum on a trait the field evidence links to cohesion is a reason that survives being written down.

Third, team profiles decay. Every exit, internal move, and hire shifts the mean, the spread, and sometimes the minimum, so a composition read from a year ago describes a team that no longer exists. Departures matter as much as arrivals: when the most agreeable member of a strained team resigns, the team's minimum moves without anyone being hired. Re-profiling at meaningful roster changes is straightforward when an assessment platform already holds current facet profiles for development purposes, and it keeps the team-side numbers as fresh as the candidate-side ones.

Selection will keep scoring individuals; the ranking, the interview, and the offer are individual instruments, and the evidence behind them is strong. What team personality composition research adds is one step at the end, while the decision is still open: recompute the destination team's statistics with the candidate included. Run it, and the two account managers hired in the same cycle stop being one decision made twice. Skip it, and the team's numbers change anyway, unobserved.

Where 5Profiler stands

5Profiler measures candidates at facet resolution, and a facet profile can be read against more than one target. Alongside role-referenced scoring, a hiring team can hold a finalist's profile up against the destination team as it stands: where the team's mean sits on the traits the work strains, how wide its spread runs, and which member currently sets its minimum. Existing members complete the same instruments a candidate does, so the team profile is measured rather than recalled. The same profiles then serve development after the hire, so the team's picture stays current as its membership changes.

Read the science behind the platform · See it on your roles

References

  1. Barrick, M. R., Stewart, G. L., Neubert, M. J., & Mount, M. K. (1998). Relating member ability and personality to work-team processes and team effectiveness. Journal of Applied Psychology, 83(3), 377–391.
  2. Bell, S. T. (2007). Deep-level composition variables as predictors of team performance: A meta-analysis. Journal of Applied Psychology, 92(3), 595–615.
  3. Humphrey, S. E., Hollenbeck, J. R., Meyer, C. J., & Ilgen, D. R. (2007). Trait configurations in self-managed teams: A conceptual examination of the use of seeding for maximizing and minimizing trait variance in teams. Journal of Applied Psychology, 92(3), 885–892.
  4. LePine, J. A. (2003). Team adaptation and postchange performance: Effects of team composition in terms of members’ cognitive ability and personality. Journal of Applied Psychology, 88(1), 27–39.
  5. Peeters, M. A. G., van Tuijl, H. F. J. M., Rutte, C. G., & Reymen, I. M. M. J. (2006). Personality and team performance: A meta-analysis. European Journal of Personality, 20(5), 377–396.
  6. Prewett, M. S., Walvoord, A. A. G., Stilson, F. R. B., Rossi, M. E., & Brannick, M. T. (2009). The team personality–team performance relationship revisited: The impact of criterion choice, pattern of workflow, and method of aggregation. Human Performance, 22(4), 273–296.
  7. Steiner, I. D. (1972). Group process and productivity. Academic Press.

© 2026 Future Proof. All rights reserved. 5Profiler™ and the 5Profiler bloom mark are trademarks of Future Proof.

See it on your roles

Evidence over intuition, on your next hire.

A 30-minute walkthrough of 5Profiler with your roles, not a canned deck — and a sample report to keep.