The strongest predictor of rated sales performance in the sales-specific research record is a form. Scored biodata, a structured questionnaire about what a candidate has actually done, predicts supervisor ratings of salespeople at .52, ahead of every instrument with more stage presence (Vinchur, Schippmann, Switzer, & Roth, 1998). Sales hiring, meanwhile, still runs as an audition: the debrief circles energy, confidence, and the pleasant sensation of having been sold to. It is the corner of hiring where folk psychology does its most expensive work, and the correction has been sitting in the field's own literature since 1998.
The folk theory is visible in any sales job advertisement: hunger, energy, natural closers, the gift of the gab under four synonyms. What the advertisement describes is presence in a room, and presence in a room is the one thing an interview measures effortlessly. The research file rewards something else: recorded persistence, the candidate who has sold before, sold measurably, and left a trail a scoring key can read.
The file in question is Vinchur, Schippmann, Switzer, and Roth (1998), a meta-analysis of predictors of job performance for salespeople published in the Journal of Applied Psychology, still the reference synthesis for hiring salespeople nearly three decades on. One design choice explains its durability: the authors refused to blend their outcome measures, estimating every predictor twice, once against supervisor ratings of performance and once against objective sales, the booked figures themselves. The two columns disagree, and the disagreement, more than any single coefficient, is what a sales hiring process can actually use.
Supervisor ratings and the sales register rank the predictors differently
A supervisor-ratings criterion measures a manager's composite impression of the year: activity visible from the corridor, polish on the calls the manager happened to join, cooperativeness, the look of effort. An objective-sales criterion measures what closed. Both are legitimate outcomes, and they are not interchangeable, because the predictors of sales performance rank one way against the first and another way against the second. Vinchur and colleagues estimated the major instrument families against each, and Figure 1 pairs the results.
General cognitive ability is the pair to sit with. Against supervisor ratings it reaches .40, stronger than any predictor in the file except the two scored instruments. Against objective sales it manages .04, indistinguishable from measuring nothing. Conscientiousness runs the other way, .25 against ratings and .41 against the register. Potency, the assertive facet, is the steady one, .28 against ratings and .26 against the register. One predictor collapses across the split, one strengthens, one holds.
The coefficients also carry their date. The 2022 re-estimates of the selection literature lowered most classical validity values, and Vinchur's corrected figures predate that re-estimation, so Figure 1 is best read as an internally consistent ranking rather than as a set of absolute magnitudes (Sackett, Zhang, Berry, & Lievens, 2022). Why the estimates moved is traced in the briefing on the general mental ability debate; where each instrument family stands in the current cross-industry table is charted in what actually predicts job performance. Neither adjustment disturbs what the pairs show: whichever era's corrections apply, ratings and the register order the predictors differently.
The .40/.04 pair deserves one more sentence of interpretation, because it indicts a criterion, not a construct. Smarter salespeople sound better in reviews, plan more legibly, and talk about territory the way managers like territory talked about, and a rating absorbs all of it; the booked number, in this file, absorbed almost none of it. A hiring process that selects sharp generalists and then wonders why the revenue distribution looks unchanged has usually made exactly this substitution, and the general case for ability measurement, which is real, belongs to different jobs and different criteria.
The dimmed row at the bottom of Figure 1 exists to make the contamination undeniable. Age correlates with supervisor ratings at .26 and with objective sales at −.06 (Vinchur et al., 1998). Using age in selection is prohibited in most jurisdictions and would be pointless on this evidence, and neither use is why the pair is shown. It is shown because a demographic attribute that tracks the ratings while tracking nothing in the register is exposing what ratings absorb: seniority read as effectiveness, tenure read as skill, a manager's picture of what a seasoned salesperson looks like. A criterion that rewards age it cannot cash in revenue will misgrade instruments too, and every ratings-column value in the file, including the flattering ones, should be read with that in mind.
Why conscientiousness sells
Of every predictor in Figure 1, conscientiousness is the only one that predicts the register better than it predicts the impression: .25 against supervisor ratings, .41 against objective sales (Vinchur et al., 1998). The direction of that gap is the diagnostic. Whatever managers watch when they rate, conscientious behavior is undersold in the watching; whatever the register counts, conscientious behavior keeps arriving in the count.
The mechanism is familiar to anyone who has run a pipeline. Most of selling happens between conversations: prospecting blocks kept when no one is checking, the fifth follow-up after the fourth silence, notes entered while the call is still warm, a forecast built on discipline instead of optimism. Those behaviors compound into the number the register books, and almost none of them photograph well in an interview.
The pattern is no sales anomaly. Barrick and Mount (1991), the meta-analysis that made the Big Five usable in selection, found conscientiousness predicting across occupational groups, sales included; the general coefficients, and the argument for treating the trait as the default personality input in hiring, are the territory of the briefing on conscientiousness at work.
At facet grain the trait splits into a striving side and a tidiness side, and selling leans hard on the first. Industriousness, the persistence-and-goal-pursuit component, is what the pipeline behaviors above describe; orderliness keeps the calendar and the records straight but places no calls. The case for measuring below the domain, and the coefficients that justify it, live in the facet-level measurement briefing. The dependability neighborhood adds one more line to the sales case: in customer-facing work, integrity-adjacent measures also forecast lower counterproductive behavior, the damage a wrong hire does between commissions (Ones, Viswesvaran, & Schmidt, 1993).
Sales hiring reads extraversion at the wrong grain
Extraversion is what sales hiring believes it is buying when it hires charisma, and the trait's pooled record does not support the belief; the full accounting, including how weak the domain-level coefficients run in sales itself, is the business of the briefing on extraversion at work. What survives that accounting is a sub-dimension. Vinchur's file carries it: potency, the assertive facet, predicts on both criteria (.28 against ratings, .26 against objective sales) — the only personality values in the file that hold their size across the split. It is also the piece of charisma that earns its weight: assertion speaks first on a cold call, asks for the close, and holds price under pushback. A candidate can carry all of that without filling a room.
What charisma mostly signals across a table is the other half, the sociable, expressive, immediately likeable side, and that half's sales record gives a hiring process nothing to stand on; the subtrait split is drawn in full in that briefing. An interview cannot help privileging it: sociability is the component that shakes hands, and it gets scored in the first minute under the label "presence." Nor does wanting the work rescue the signal. Enthusiasm for selling is a direction, and the briefing on interests versus ability measures how little direction adds to a performance forecast.
The correction is a change of grain, and it costs nothing but resolution. Read at domain level, extraversion mostly reports how the candidate interviews. Read at facet level, it contains one component with a defensible, modest claim on a sales profile and several with none. A screening profile that weights assertiveness where the role genuinely runs on initiative, and leaves gregariousness unweighted, is using the trait the way its evidence permits, which is narrower than the job advertisement assumed and considerably more useful. An instrument that reports extraversion as a single score cannot support that weighting at all; the distinction it needs was averaged away before the profile was printed.
Empirical keying turns biography into an instrument
The top of the ratings column belongs to instruments nobody demos on a conference stage. Scored biodata predicts supervisor ratings at .52, and structured sales-ability inventories predict them at .45, both ratings-criterion estimates, both ahead of cognitive ability and every trait measure in the file (Vinchur et al., 1998). That both values sit in the ratings column matters as much as their size: the file's two strongest instruments are established against the criterion the dimmed age pair cautioned about. The caution reaches the coefficients and leaves the method intact, since a scored biography and a scored inventory both sample recorded behavior, which is the material a register counts too.
Biodata is biographical fact collected as data: which jobs, for how long, what was sold, what was carried as a target, what was done outside work that looks like initiative. What converts the biography into an instrument is the key, built and cross-checked on outcome data before anyone is scored with it. A résumé contains much of that raw material and none of the scoring; a reader supplies the weights from memory and mood. The line between reading biography and scoring it is the subject of the résumé-screening briefing, and it separates one of hiring's weakest screens from one of its stronger instruments. Nor is the showing a sales quirk: biodata performs across the broader selection literature as well (Schmidt & Hunter, 1998).
Sales-ability inventories work the adjacent seam. They are structured item sets about selling situations and the working knowledge of the craft, scored against a key, asking directly the question the charisma interview asks by proxy: does this person know how selling is done. The advantage of both instruments is procedural before it is psychological. Every applicant answers identical questions, every answer meets an identical key, and the key answers to outcomes; an interviewer's memory of a biography does none of this, and ten confident minutes can rewrite it.
A key, though, is only as good as the outcome it was built on, which puts the criterion problem back on the table at the level of a single hire. Quota attainment at month twelve, the standard cell in the standard dashboard, is a snapshot, and hires who look identical in it can have contributed very different years. Figure 2 sketches the shapes the snapshot hides: a fast ramp, a typical one, and the ramp that never arrives. The gaps between them, months of quota coverage, pipeline carried early, the seat that empties before the year does, are outcomes a scoring key can be built against and a charisma read cannot.
A battery built for the register, then verified in the room
Assembled, the evidence yields a sales assessment with five working parts, drawn in Figure 3. The parts are not novel. The discipline is the assembly: each instrument is present for a criterion it demonstrably predicts, and the weights do the adapting across roles while the components stay put.
The objection-scenario judgment test samples the moment charisma is supposed to be for: a prospect pushes back on price, a champion goes silent, a discount is demanded on the last day of the quarter. Situational judgment items put those moments in front of every candidate identically and score the responses against a key, so what a panel once inferred from confidence, the battery samples directly, before anyone's likeability has entered the record.
The interview comes last and changes function. With scored biodata, a facet profile, and scenario judgment already in hand, panel minutes stop hunting for personality across a table and go to verification: probing the outlier claims, resolving what the instruments left ambiguous, hearing the craft knowledge no key reaches. The mechanics of scoring that conversation, anchored questions, independent ratings before discussion, are the subject of the briefing on running structured interviews.
The weights are where the battery meets the role. A new-logo pursuit desk, cold outreach, short cycles, constant refusal, plausibly leans the profile toward industriousness and the assertive facets; an account-growth desk, long relationships, renewal economics, service recovery, leans it toward follow-through and steadiness. Those are role-analysis judgments, and the point is to make them explicit and record them, the way a role profile on the platform carries them; a sample report shows the facet grain the weighting needs to be legible at all.
What the weights cannot do is rescue a battery aimed at the wrong outcome. Every component in Figure 3 is scored against something, that something has to be chosen before the first candidate answers anything, and a key built on rated cooperativeness will select for rated cooperativeness with perfect fidelity while the revenue plan goes unserved. The assembly question therefore comes second. The criterion question comes first, and it is the one most sales organizations never formally answer.
Choosing the criterion is the decision sales hiring skips
The professional standards for selection state the sequence without ceremony: define the outcome the job is supposed to produce, then choose instruments on the evidence that they predict it (Society for Industrial and Organizational Psychology, 2018). Sales inverts that order more readily than most functions, because the work looks as though it arrives with its criterion attached. It does not. Booked revenue, gross margin, retained accounts, new logos, forecast accuracy, and a manager's read of how a territory is being run are separate outcomes, and the paired bars in Figure 1 show that picking among them changes who should be hired. A team hiring for the register and a team hiring for the rated year can read one candidate file, reach opposite decisions, and both be defensible.
The weighting then follows from the choice instead of from taste. Where the outcome is revenue, the unglamorous entries carry the load: scored biography, industriousness, documented pipeline behavior that survived a bad quarter, with conscientiousness sitting at .41 against objective sales as the strongest personality claim in the file. Where the outcome is rated performance, cognitive ability at .40 and the sales-ability inventory at .45 against that criterion have a genuine claim, provided the process holds on to what the age row exposed about everything a rating touches. Most sales organizations want some of both, which makes the weights a compromise that someone should set deliberately and sign, before a debrief settles the matter by accumulation.
Charisma keeps a place in either version and a smaller place than it now occupies. The defensible size of that place is potency's moderate, two-criterion showing: real, worth measuring, and nowhere near the share of the decision that one conversation across a table currently awards it. The rest of what reads as charisma across a table has no comparable claim on a sales profile, so a screening key that scores assertiveness and leaves sociability unweighted is being neither austere nor squeamish about personality. It is spending its weight where the evidence sits.
The outcome also has to be recorded in a shape a scoring key can consume, which is where most sales hiring programs stop. Quota attainment at month twelve is a checkbox, every trajectory in Figure 2 either clears it or misses it while saying nothing about the year that produced the result. Months to first close, months to full coverage, pipeline built in the opening quarter, and the exit date where there is one are the fields that make a sales assessment improvable, because each of them can be joined back to what the battery said at the door. Organizations that hire at volume keep records in roughly this form because the funnel gives them no choice; the operating pattern for measuring outcomes at that scale is worked through in the briefing on high-volume hourly hiring, and a sales function hiring at a small fraction of that volume can borrow the pattern intact.
Most organizations find, at the point of attempting this, that the only performance data they hold in usable shape is a manager rating. That is a reporting problem with a known fix, and it is worth fixing before the next hiring cycle instead of after it. The revenue systems already carry the register; what they typically lack is a durable join between a person, a start date, and a monthly number, which is an afternoon of data engineering rather than a research programme. Until the join exists, a sales hiring process can still act on this evidence by leaning on the predictors that hold their size across both criteria, conscientiousness and potency, and by treating every ratings-only estimate as provisional. The weights get sharper the moment the register becomes readable at the level of an individual hire.
One habit has to go for any of this to hold. When the panel meets before the measured file exists, the impression becomes the file, and every instrument that arrives afterward is read for confirmation of a decision that has already been made. Running biography, profile, and scenario judgment ahead of the conversation costs nothing but scheduling, and it is the condition under which the interview can do the thing interviews do well.
The audition will survive this evidence, because meeting a talented salesperson is a pleasure and reading a scoring key is not. The record makes one thing hard to defend: letting the pleasure set the weights. A sales organization that names its criterion, scores biography and behavior against it, and reserves the room for verification will still hire people who are a delight to meet. It will simply stop mistaking the delight for the forecast.
Where 5Profiler stands
5Profiler applies role-referenced scoring at facet grain, which is what the sales profile this evidence recommends actually requires: industriousness and assertiveness weighted for the desk being hired, gregariousness reported and weighted only as far as its record supports. A new-logo pursuit desk and an account-growth desk are configured as separate role references, so two sales teams hire on different weights from one battery. Scenario items and structured interview guides hang off that role definition, which keeps panel time pointed at verification. The weighting a role carries is printed on the candidate's report, so a hiring manager can see what the score was tuned for.
References
- Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26.
- Ones, D. S., Viswesvaran, C., & Schmidt, F. L. (1993). Comprehensive meta-analysis of integrity test validities. Journal of Applied Psychology, 78(4), 679–703.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.
- Society for Industrial and Organizational Psychology. (2018). Principles for the validation and use of personnel selection procedures (5th ed.). Industrial and Organizational Psychology, 11(Suppl. 1), 1–97.
- Vinchur, A. J., Schippmann, J. S., Switzer, F. S., III, & Roth, P. L. (1998). A meta-analytic review of predictors of job performance for salespeople. Journal of Applied Psychology, 83(4), 586–597.