Every famous construct in hiring can be given the same audit: set the claim that made it famous beside the pooled record of the studies that tested it. Emotional intelligence hiring is the audit's canonical case. The claim, published in the Harvard Business Review in 1998, was that emotional intelligence had proved twice as important as IQ and technical skill for excellent performance. The pooled record, assembled over the decade that followed, put the ability model's validity for job performance at .18, and found the popular questionnaires that carry the label to be a blend of traits psychology already had.
The gap is not a scandal, and this briefing does not treat it as one. It is the visible end of a cycle that runs through applied psychology on a schedule: a compelling label, then a commercial claim, then, years later, a meta-analysis that finds mostly old traits in new packaging. Emotional intelligence ran the cycle in full. Growth mindset ran it at smaller scale. Grit ran it too, and this series has already published the result.
The cycle matters because organizations adopt during its middle years: the interview loop gets an empathy score, the leadership model gets an EQ pillar, the assessment budget shifts toward branded questionnaires. Each of those decisions sits downstream of a sentence, and the sentence is checkable. This briefing runs the check.
The cycle has a shape: label, claim, meta-analysis, residue
A construct becomes fashionable in a fixed sequence. First the label: a researcher or author names something managers already believe in, and the label works because the observation underneath it is real. Some people do handle people better than others; some people do persist longer. Then the claim: a quantitative-sounding statement of how much the new construct matters, which travels through business media at a speed no validation study can match. Then a long silence while a literature accumulates. Finally the meta-analysis, which asks the questions the claim never faced: how much does the construct predict, and how much of it is new?
The residue is the recurring answer to that second question. When popular psychology constructs are decomposed, what remains is usually something psychology already had: most often a trait or two from the established Big Five taxonomy, whose structure and workplace record this series covers in its own briefing, sometimes joined by general cognitive ability. The new label is rarely empty. It is usually full of something old.
Hiring hosts this cycle unusually well. The problems are real, so a label that names one meets little resistance. The constructs are plausible; nobody who has managed a team doubts that self-regulation matters. And the claim is unverifiable at the moment of adoption, because the evidence that could verify it does not exist yet. Meta-analysis needs a literature, and the literature takes a decade. An organization evaluating a new construct in year two of the cycle is being asked to accept the label's promise on a schedule set by the vendor, while the evidence is still years from existing.
No stage of the sequence requires bad faith. Labels get coined because researchers notice real regularities; claims travel because business media reward a strong number; the correction arrives late because pooling requires studies that take years to accumulate. The audit exists precisely because good faith is no safeguard. A sincere claim that outruns its evidence misdirects a hiring process exactly as a cynical one does.
Emotional intelligence hiring rests on a split construct
Any statement about EQ test validity has to begin with a distinction, because the label covers two research traditions that measure different things. The ability model, whose principles Mayer, Caruso, and Salovey (2016) set down formally, treats emotional intelligence the way psychology treats any intelligence: as a capacity to reason accurately about emotions and to use emotional knowledge to enhance thought. Ability EI is measured with tasks that have better and worse answers, such as identifying the emotion in a face or judging which regulation strategy fits a situation.
Mixed models are the other tradition, and the one most commercial EQ instruments descend from. They blend self-rated effectiveness across a wide territory: optimism, assertiveness, self-regard, stress tolerance, drive. The measurement is questionnaire self-report, which means a mixed-EI score records how emotionally effective a person says they are.
The distance between asking and testing is itself measurable. Joseph and Newman (2010) report that self-reported ability EI and performance-based ability EI, the same construct asked about versus tested, correlate just .12. Self-assessment and measurement are nearly strangers even inside the ability model, before the mixed models widen the definition further.
The distinction is not academic bookkeeping; it decides what an emotional intelligence hiring program is actually measuring. An organization that adds an ability-EI task to its process has added a narrow cognitive measure with a context-dependent record. An organization that adds a mixed-model questionnaire has added a self-description instrument whose contents need a decomposition before the score means anything. The two additions look identical on a product sheet, and they are two different decisions.
The evidence arrived a decade later, and it was modest
The claim predates both traditions' mature literatures. In 1998, Daniel Goleman, writing in the Harvard Business Review, reported what he had found in proprietary competency models from 188 companies:
when I calculated the ratio of technical skills, IQ, and emotional intelligence as ingredients of excellent performance, emotional intelligence proved to be twice as important as the others for jobs at all levels.Daniel Goleman, Harvard Business Review, 1998
The article went further at the top of the organization, attributing nearly 90% of the difference between star and average senior performers to emotional intelligence factors (Goleman, 1998). The sentences were not offered as provocation. They were offered as findings, from models built inside real companies, and that is what made them consequential — and, eventually, auditable.
Competency models are not trivial sources; they encode what organizations believed distinguished their best people. But belief encoded is still belief, and the ratio was a claim about prediction, checkable against studies that measure rather than remember.
The pooled test came in 2010. Joseph and Newman published an integrative meta-analysis of the emotional intelligence literature in the Journal of Applied Psychology, pooling the available studies linking EI measures to job performance. For the ability model, the kind measured with right-and-wrong-answer tasks, the pooled validity was .18 (k = 10).
That pooled figure is not evenly distributed across jobs. How much ability EI predicts depends on how much of the work consists of managing one's own and others' feelings, the demand researchers call emotional labor (nursing, collections, front-line service). Figure 1 shows the split, with the 1998 sentence alongside for scale, and the implications section returns to what it instructs a hiring team to do.
Joseph and Newman's cascading model explains the dependence: emotion perception feeds emotion understanding, understanding feeds emotion regulation, and regulation is the step that reaches performance. In a role where regulating feeling is the work, the cascade has somewhere to land. Where it is not the work, the chain ends on nothing the job requires.
A .18 that rises to .24 in emotionally demanding work is a real finding, and a usable one. It also rests on a thin base: ten usable studies, two decades into the construct's fame, a count that is itself information about where the field's attention went. What did not survive pooling in any form was the shape of the original claim: a general ingredient, dominant at every level, in every kind of job. The record that exists is modest, conditional, and specific.
Seven established constructs reproduce the mixed-model score
Almost none of that record describes the instruments most organizations actually encounter. Those are the mixed-model questionnaires, and pooling them produced the finding that decides the case.
Mixed-model EI looked stronger in that 2010 meta-analysis: .47 against job performance overall, and, unlike the ability model, it predicted everywhere, at .59 in high emotional-labor jobs and .43 in low. That result should have raised suspicion on its face. A self-report questionnaire outpredicting the ability version of its own construct, in contexts where the construct should not matter, is a puzzle with two candidate solutions: either self-rated emotional effectiveness is a remarkable new signal, or the questionnaire is measuring several old signals at once.
Joseph, Jin, Newman, and O'Boyle (2015) resolved the puzzle in the Journal of Applied Psychology, in a paper whose title poses the question directly: why does self-reported emotional intelligence predict job performance? Their updated estimate put mixed EI at .29 against supervisor-rated performance. Then they decomposed it. A weighted composite of seven established constructs — performance-based ability EI, self-efficacy, self-rated performance, conscientiousness, emotional stability, extraversion, and general mental ability — reproduced mixed-EI scores at a multiple R of .79, the correlation between the best-weighted blend of the set and the questionnaire's own scores.
The control condition delivered the verdict. With that set held constant, mixed EI's relationship to job performance dropped to nil: β = −.02. The questionnaire's validity was real, and it was borrowed. Conscientious, emotionally stable, extraverted people who rate their own effectiveness highly score high on mixed EI, and those characteristics, each measurable under its own name, do the predicting. Emotional stability's workplace record is documented in its own briefing in this series. The ingredients are well-measured constructs with records of their own, and mixed EI's contribution is to relabel them. Figure 2 maps the decomposition.
The ability model survives that test, barely. Entered after the Big Five and cognitive ability, performance-based ability EI adds 0.2% of explained variance to the prediction of job performance (ΔR² = 0.2%; Joseph & Newman, 2010). Ability EI measures something of its own, so distinctness is not the difficulty. Increment is: the traits and the ability score a standard battery collects first already anticipate almost all of it.
This is the audit's general lesson. A correlation with performance is the beginning of a validity case, not the end of one; instruments are made of constructs, and until the constructs are named and controlled, a validity coefficient cannot say whether an organization is being offered new information or measures it already runs under another name.
Growth mindset ran the cycle at smaller scale
Growth mindset, the belief that ability is malleable, moved from education research into corporate values statements and interview rubrics on the strength of a similar promise: that the belief itself drives achievement, and that instilling it changes outcomes. Growth mindset hiring, in which interviewers listen for mindset language and score it, is the construct's selection-stage form. The pooled record comes from the construct's home domain, education, where the studies actually exist.
Sisk, Burgoyne, Sun, Butler, and Macnamara (2018) synthesized that record in Psychological Science. Across 273 effect sizes covering 365,915 people, mindset correlates with academic achievement at roughly .10, about 1% of the variance in achievement. Their second meta-analysis, of studies that tried to install the belief rather than merely measure it, found a small positive average effect on achievement.
Read fairly, the interventions are not empty. Benefits were somewhat larger for academically at-risk and lower-income students, which suggests a targeted educational tool with a defensible use. What the record does not suggest is a trait-like lever on adult workplace outcomes.
The record is also a record of school achievement, mostly in students, in the domain where the theory was built and the interventions were designed. Selection is a harder test. A hiring use imports the construct into a new population, a new criterion, and an adversarial setting in which candidates know what the interviewer wants to hear, and it does so with no pooled evidence of its own. The education numbers are the best case currently on offer.
The follow-up made the estimate smaller still. Macnamara and Burgoyne (2023) re-synthesized the intervention literature in Psychological Bulletin, across 63 studies and 97,672 participants, and put the average effect at d = 0.05, a standardized mean difference that was nonsignificant once publication bias was corrected. Figure 3 plots the record at both dates, with the confidence intervals and the highest-quality subset. A construct can be humane in spirit and useful in a targeted classroom setting while contributing nothing measurable to a hiring decision, and on the current evidence mindset is that construct.
Grit is the pattern's third instance, and its audit has already run. Introduced by Duckworth, Peterson, Matthews, and Kelly (2007) as perseverance and passion for long-term goals, grit gets its full accounting in the conscientiousness briefing, which reports the value that settles most grit assessment decisions: pooled across the literature, the higher-order grit construct tracks conscientiousness at about .84 (Credé, Tynan, & Harms, 2017).
Three questions to ask before the next label arrives
The pattern repeats reliably enough to be run as a standing audit, one that needs no statistician on staff, only the published literature and an afternoon. Each question below has a demonstration already on the table.
- What does it correlate with among the constructs you already measure? A genuinely new construct correlates modestly with the old ones. Mixed EI failed the question completely: a weighted blend of established measures reproduced its scores almost in full. Any instrument vendor with validation data can compute this; a vendor who has not is asking for adoption on the label alone.
- Does its validity survive controlling the established constructs? This is incremental validity, the prediction a measure adds after the ones you already have are entered first. Mixed EI's validity fell to effectively nothing under that test, and tested ability EI kept its distinctness while adding a fraction of a percent of explained variance.
- Who verified the headline claim, and with what data? The sentence that anchored two decades of emotional intelligence hiring came from proprietary competency models no independent researcher could reanalyze. The meta-analyses that later contradicted it were built from published studies anyone can re-pool. A claim's checkability is part of its evidence, and a buyer who cannot obtain the underlying data has learned something important about the claim.
The audit also has a null form: when no pooled record exists at all, that absence is the finding. An organization adopting an EI battery in 1999 had nothing to check the claim against, and that, read correctly, was the answer to how much weight the claim could bear. Where a decision cannot wait for a pooled record, the defensible default is to measure the established constructs the new label most resembles, and to revisit when the record exists.
Measure the established taxonomy well, and admit new labels on increment
The constructive reading of these audits is that popular labels keep pointing at real territory. The material emotional intelligence names (composure under interpersonal load, accurate perception of others, self-regulation) is measurable, but as facets of the established taxonomy, each answering to a record of its own. Other briefings work the point on their own traits: openness carries the learning-and-curiosity record that innovation programs reach for, and the case for measuring traits at facet grain is made in the facet-level briefing. A hiring organization that measures the taxonomy well already holds most of what the labels promise. The reverse is also true: an organization that swaps its established measures for each season's construct is trading documented records for undocumented ones, repeatedly.
Treat any claim of the form "X matters more than IQ" as a proposition about a regression. A more-important-than claim implies that someone entered both constructs into one model against the same criterion and compared their weights, which implies a decomposition table. Ask to see it. For emotional intelligence, when that table was finally assembled from pooled public data, the comparison reversed: the constructs said to be outweighed did the predicting. The demand applies prospectively too. A vendor who publishes decomposition tables against the Big Five and cognitive ability makes the audit quick; a vendor who publishes only a standalone validity coefficient, nothing controlled, is presenting the mixed-EI situation as it stood before Joseph and colleagues took it apart.
The emotional-labor result is where this evidence becomes actionable. Validity rises to .24 in jobs high in emotional labor and falls away entirely in work that makes no such demand, turning negative there once the Big Five and cognitive ability are entered. Read as an instruction, that is a finding about roles rather than about instruments, and it is more useful to a hiring team than the overall coefficient.
The translation is therefore role analysis. Name the emotional demands the work actually carries, then measure the traits and abilities that meet them through a battery built from established instruments. Where those demands are absent, a total score labeled EQ adds a name, not information, over the constructs it is made of.
Incremental validity is the standing entry rule all of this implies. A new construct joins the battery when it demonstrates prediction beyond the established traits plus cognitive ability, on data someone outside the vendor can inspect. That gate is not hostile to novelty; it is how genuine novelty gets recognized. Everything currently inside a well-built battery passed through it. Holding the gate also keeps the battery short: every redundant construct spends candidate minutes that a genuinely distinct measure could have used.
The cycle will run again. The next construct will name something real, attach a number, and spend its first decade unpooled. When it reaches your shortlist, run the audit: set the famous claim beside whatever pooled record exists, ask what the construct correlates with, ask what it adds, and read the absence of a record as an answer too. The constructs that survive that meeting have, so far, been the old ones, measured well.
Where 5Profiler stands
5Profiler measures the constructs that keep turning up underneath the fashionable labels: the Big Five at 30-facet resolution, with every scale reported under its published name and answerable to its published record. That construction makes the comparison this article describes a practical exercise, because the established measures are already in the candidate's profile, so any construct proposed as an addition has something concrete to demonstrate increment against, instead of a label taken on trust. Where a role carries genuine emotional demands, facet-level trait measurement addresses those demands directly, scale by scale, with no branded total in between.
References
- Credé, M., Tynan, M. C., & Harms, P. D. (2017). Much ado about grit: A meta-analytic synthesis of the grit literature. Journal of Personality and Social Psychology, 113(3), 492–511.
- Duckworth, A. L., Peterson, C., Matthews, M. D., & Kelly, D. R. (2007). Grit: Perseverance and passion for long-term goals. Journal of Personality and Social Psychology, 92(6), 1087–1101.
- Goleman, D. (1998). What makes a leader? Harvard Business Review, 76(6), 93–102.
- Joseph, D. L., Jin, J., Newman, D. A., & O’Boyle, E. H. (2015). Why does self-reported emotional intelligence predict job performance? A meta-analytic investigation of mixed EI. Journal of Applied Psychology, 100(2), 298–342.
- Joseph, D. L., & Newman, D. A. (2010). Emotional intelligence: An integrative meta-analysis and cascading model. Journal of Applied Psychology, 95(1), 54–78.
- Macnamara, B. N., & Burgoyne, A. P. (2023). Do growth mindset interventions impact students’ academic achievement? A systematic review and meta-analysis with recommendations for best practices. Psychological Bulletin, 149(3–4), 133–173.
- Mayer, J. D., Caruso, D. R., & Salovey, P. (2016). The ability model of emotional intelligence: Principles and updates. Emotion Review, 8(4), 290–300.
- Sisk, V. F., Burgoyne, A. P., Sun, J., Butler, J. L., & Macnamara, B. N. (2018). To what extent and under which circumstances are growth mind-sets important to academic achievement? Two meta-analyses. Psychological Science, 29(4), 549–571.