Every selection method makes two promises: that it predicts, and that it predicts for the reason printed on the label. Most methods that fail break the first promise. The assessment center broke only the second. Run candidates through a day of simulated work (an in-basket of competing decisions, a leaderless group discussion, a role-played confrontation with a struggling direct report), have trained assessors rate every performance on dimensions like planning, influence, and judgment, integrate the ratings, and the resulting score genuinely forecasts later success. Look inside the ratings, though, and the dimensions on the scorecard are missing. The method works; the explanation collapsed under its own data.
The field admitted this in print a generation ago. Klimoski and Brickner (1987) gave their Personnel Psychology review a title that asked, without irony, why assessment centers work, and they meant the question literally: the evidence that centers predict managerial success was strong, and the understanding of why was not. The flattering story, stable leadership traits captured in the exercises and expressed later on the job, had to compete with alternatives they took seriously, among them the possibility that people are simply consistent performers across settings, and the possibility that assessors and promotion committees reward the same polish twice, contaminating the criterion. The decades since have been unkind to the flattering story and kind to the method.
The inversion is not trivia, because an assessment center is a heavy commitment: days of candidate and assessor time, purpose-built simulations, trained raters, a consensus meeting at the end. What the organization believes all that effort delivers is the competency taxonomy, a set of labeled windows into a candidate's leadership. What the evidence says it delivers is a set of well-designed exercises, observed carefully and summed. The difference decides what the reports may claim, how much weight the profile pages deserve, and whether a simpler instrument would capture the same signal.
The scorecard assumes dimensions travel
The assessment centre method, in the British spelling much of its literature still carries, has a stable architecture wherever it runs. Candidates rotate through several exercises built to simulate the demands of a target role. Multiple trained assessors observe, take behavioral notes, and rate each candidate on a fixed set of dimensions after each exercise. The ratings are then integrated, in a consensus discussion or by arithmetic, into dimension scores and an overall assessment rating. In its leadership assessment center form, the method decides promotions and development investments as often as external hires.
Underneath the logistics sits a data model, and it is worth making explicit because everything that follows is a claim about its shape. The design is a matrix: exercises down the side, dimensions across the top, one rating in each cell. The left panel of Figure 1 shows the assumption the matrix encodes. A dimension is supposed to be a property of the person, so ratings of the same dimension should travel together across exercises: the candidate who shows strong influence in the group discussion should show strong influence in the negotiation. If that holds, the columns cohere, the dimension profile is a measurement, and feedback like "strong on planning, weaker on influence" describes the candidate, not the afternoon.
A great deal rides on that assumption, because the dimension profile is the output organizations lean on hardest. Development plans are written from it, coaching engagements are scoped against it, and succession files carry it forward for years. The overall rating decides who advances; the profile claims to explain why, and that explanation is much of what sets the method apart from a test score. If the columns of the matrix do not cohere, the most heavily used output is the least supported one.
The factors that emerged were exercises, not dimensions
The assumption is testable, and Sackett and Dreher (1982) tested it in the Journal of Applied Psychology with rating data from three organizations and 559 candidates. Factor analysis, a statistical technique that finds which ratings rise and fall together, should have recovered clusters shaped like the scorecard's columns. It recovered the rows. The factor patterns represented exercises rather than dimensions: ratings of supposedly distinct dimensions made in the same exercise moved together, while ratings of the same dimension made in different exercises barely moved together at all. In two of the three organizations, the mean correlation among across-exercise ratings of the same dimension was near zero.
Translate that into one candidate's file. Her planning rating from the in-basket agreed with her other in-basket ratings far more than it agreed with her planning rating an hour later in the group exercise. The instrument was scoring performances, not properties. The right panel of Figure 1 is the structure the data keeps returning: coherence runs along exercises, and the dimension columns dissolve.
The result was not an artifact of early scoring habits. Wirz and colleagues (2020) took a harder version of the question to three modern samples. Internal structure can always be argued about, so they ran the external test, which is difficult to argue with: if overall dimension ratings measure real constructs, they should line up with measurements of the same constructs taken elsewhere, by other methods and other raters. Their factor analyses found source factors, groupings by where a rating came from, and no dimension factors at all. Surveying the accumulated record, Lance (2008) drew the conclusion the field had long resisted: post-exercise ratings substantially reflect the exercises that produced them, behavior is situationally specific, and centers should be designed around tasks and roles rather than around dimensions. The strongest opposing view is on record: Melchers and König (2008) argue it is not yet time to dismiss dimensions. The mainstream verdict, though, has held for decades, and it is negative.
There is nothing mysterious about why exercises dominate. A leaderless group rewards reading a room of peers in real time; an in-basket rewards solitary triage under time pressure; a role-play turns on composure in one difficult conversation. These are different situations making different demands, and candidates respond to situations. Read that way, the finding is less a criticism of the method than a fact about people that the dimension scaffolding was built to overlook.
Aggregation recovers a usable overall score
If exercise factors dominate the ratings, where does the predictive power live? Kuncel and Sackett (2014) supplied the resolution in the Journal of Applied Psychology, and it reframed the whole dispute. Their argument is about what happens at the aggregation step, the moment a candidate's post-exercise ratings are combined into overall dimension scores. Combine enough of them and exercise-specific variance recedes, because each exercise's quirks are diluted by the others, exactly as measurement theory says they should be.
That sounds like vindication for the scorecard until you ask what the surviving variance is made of. When Kuncel and Sackett decomposed the aggregated dimension scores, the largest source of dimension variance was a general performance factor common to every dimension, not the distinct constructs printed on the report. The final scorecard mostly carries one signal under several labels. An overall assessment rating, or any mechanical composite of the dimension scores, stands on defensible ground: it pools many observations and lets their noise cancel. What loses its footing is profile interpretation, the practice of reading meaning into the gap between a candidate's judgment score and her influence score, because those two numbers are largely the same general signal plus differently flavored error.
Assessors should "regularly sum multiple measures of the same construct to both reduce error as well as accumulate shared construct relevant variance."Kuncel & Sackett, 2014
Sum, in other words, and resist interpreting the parts. The logic is the one that makes a long test more reliable than any single item: no competent psychometrician interprets a lone item, and a lone exercise rating deserves the same humility. The advice lands on the same spot as Lance's, from the opposite direction: he would redesign the center around tasks and drop the dimension pretense, they would keep the ratings and trust only their aggregate. Either way, the unit that means something is the whole, not the profile.
The result also redraws the integration meeting's job description. If the profile's peaks and valleys are mostly error around one general signal, a panel spending an afternoon debating whether a candidate's judgment truly exceeds her influence is deliberating over noise. The defensible use of that afternoon is quality control on the observations themselves: did the assessors see enough behavior, were the anchors applied as written, and does an outlier rating trace back to something one observer actually saw.
Assessment center validity survived the autopsy
The construct autopsy would be a curiosity if the predictions were weak. They are not. Gaugler, Rosenthal, Thornton, and Bentson (1987) pooled 107 validity coefficients from 50 studies in the classic meta-analysis of the method and estimated a corrected mean validity of .37 for the overall rating. Their internal findings still shape how the number should be read. Validity was higher when the criterion was ratings of potential rather than measures of performance, a pattern that sits uncomfortably close to Klimoski and Brickner's contamination alternative, since the polish that impresses assessors may be the same polish that impresses whoever rates potential. And in the subset of studies predicting job performance, the mean observed correlation was .25 (44 studies, N = 4,180, as re-tabulated by Sackett et al., 2022).
The modern revision trimmed the headline. Sackett, Zhang, Berry, and Lievens (2022) judged the classic range-restriction correction not credible, put the estimate at .32 with the best available correction, and published .29 (SD of ρ = .09) as their headline value under less aggressive range-restriction corrections; the statistical case behind that series-wide revision is examined in the general mental ability article. On the revised table, the assessment center sits below the structured interview (.42) and roughly beside general mental ability (.31); the full ranking is charted in the evidence review that opens this series. A validity of .29 is far from nothing. It is also far from what the method's footprint implies.
The revised estimate carries a second value worth reading alongside the first: that standard deviation of ρ of .09 says validity varies meaningfully from one operation to the next. Some centers beat .29 and some fall well short, and the construct record suggests where to look for the difference: in whether the exercises actually sample the target role's demands, whether assessors are trained against behavioral anchors, and whether the scores that reach decisions are aggregates or fragments. The number is an average over a practice run unevenly, not a property of the format.
Dimension-level evidence tells the same story from another angle. Arthur, Day, McNelly, and Edens (2003) collapsed the 168 dimension labels in use across the literature into six and estimated criterion-related validity for each. The result, plotted in Figure 2, is a narrow band running from .25 for consideration and awareness of others to .39 for problem solving. A regression composite of four dimensions reached a multiple R of .45, against 14% of criterion variance for the overall-assessment-rating benchmark. That is another win for summing over judging: a weighted combination of the ratings outpredicts the consensus number negotiated in the integration meeting. And the band's narrowness is what a general factor predicts. If the dimensions mostly carry one signal, none of them should predict dramatically better than the rest, and none does.
There is a practical reading of that band for anyone tempted to judge a center by its taxonomy. The dimensions do predict, but they predict as a bundle whose members are hard to tell apart, much as differently labeled draws on one general signal would. Choosing a center because its model features more competencies, or because its labels mirror your in-house leadership framework, is choosing between wrappers. The composite is where the validity lives, and the composite does not care what the columns are called.
The most expensive method buys mid-table validity
All of this would stay academic if the method were light to run. It is among the most resource-intensive selection procedures per candidate that organizations still use: exercises to design and stage, assessors to train and pull away from their day jobs, hours of observation per person, and an integration meeting to close it out (Thornton & Rupp, 2006). The footprint is easy to underestimate because most of it is time: senior managers serving as assessors for days at a stretch, candidates pulled out of productive work, and a long tail of report-writing and feedback meetings, recurring for every cohort assessed.
That intensity would be defensible if it secured validity nothing simpler could reach. It does not. Lievens and Thornton (2005) state the comparison plainly: assessment center validity is not higher than that of less expensive predictors such as highly structured interviews or situational judgment tests. Figure 3 puts those three methods on the two axes the decision actually turns on, operational validity against resource intensity per candidate. The structured interview holds the position that should command attention: validity of .42 at moderate intensity, above the center at a fraction of its footprint. The center itself occupies the opposite corner, the most demanding position on the chart for a validity of .29 that sits in the middle of the band; the situational judgment test is plotted alongside them at the light end of the axis. The chart carries no dollar axis because the literature offers none, but the ordering itself is not in dispute.
So what does all that intensity actually deliver? The parts of the method the construct literature vindicates are observational: several chances to watch actual behavior, trained observers, behaviorally anchored rating scales, and mechanical aggregation at the end. That bundle is portable. Structured interviews package the same observing into an hour, with anchored scales, multiple raters, and mechanical combination. Work simulations carry the realistic-task ingredient with far less of a candidate's and an organization's time; the work sample article maps where high fidelity helps and where it stops helping. Each of those parts can be obtained separately and validated separately. The center is best understood as a bundle of them plus a taxonomy the data cannot find.
Keep the discipline, not the taxonomy
The evidence converts into different instructions for the people who run centers and the people weighing whether to adopt one. For those running them, the changes follow directly from the studies above.
- Design around the work, not the wordlist. Lance (2008) draws the design conclusion from the construct record: if behavior is situationally specific, choose exercises because they sample the role's actual demands, and accept that what you measure is performance in those situations.
- Aggregate before anything reaches a decision. A dimension score from a single exercise is mostly exercise. Sum across exercises and assessors, per Kuncel and Sackett (2014), and let no fragment travel alone.
- Validate the overall rating, and claim only what it shows. The defensible product of a center is the aggregate score; track it against later performance in your own population, because the published validities are averages and your roles are specific.
- Demote the profile to conversation. If dimension profiles reach candidates at all, present them as openers for a development discussion, not as measurements. The construct evidence will not carry more weight than that, and reports that pretend otherwise borrow credibility the data has already refused.
For those deciding, the test is comparative. Ask any provider for validity evidence about the overall rating, not testimonials about the dimension reports, and then price the center against a structured interview program built with the same rigor and against structured multi-method testing of the kind described in our platform overview; the criteria in our guide to comparing assessment platforms apply here unchanged. If the center survives that comparison for a given role, it will be because its exercises sample the role's work in ways nothing lighter can, not because its competency taxonomy is finer-grained. Whichever route is taken, keep the scores and check them against performance later; a provider that cannot show a validity record for its overall rating is asking to be graded on reputation.
Centers used for development rather than selection face a narrower version of the same verdict. Watching managers work through a realistic scenario and feeding the observations back is a defensible learning design, and nothing here argues against it. What the construct evidence removes is one specific claim: that the dimension profile locates a manager's underlying strengths and gaps with measurement-grade precision. A development conversation can absorb that news and continue. A succession model keyed to profile shapes cannot.
The method will keep its place in leadership pipelines, and the evidence grants it a real one: it predicts, and the revised estimates say so without romance. Run well, it also produces something a test score cannot: a shared, behaviorally anchored record of what each candidate actually did. What was never real is the promise on the scorecard, six clean windows into one person's character. What predicts is older and plainer: watch people do the work, more than once, with trained eyes and anchored scales, then add it up. The labels made the watching look like measurement. The watching was the measurement.
Where 5Profiler stands
5Profiler delivers structured behavioral observation with anchored scoring inside the same assessment flow as its other instruments, with no rented suite or assessor roster required. Candidates work through role-relevant structured tasks, and every rating lands against a behavioral anchor. Several observations accumulate for every candidate before any score reaches a report. Where the evidence supports a measurement, the report claims one; where a dimension profile can support only a development conversation, the report frames it as exactly that, and decisions ride on the aggregate. Scores are combined mechanically, never negotiated in an integration meeting, with role-referenced scoring as the final layer.
References
- Arthur, W., Jr., Day, E. A., McNelly, T. L., & Edens, P. S. (2003). A meta-analysis of the criterion-related validity of assessment center dimensions. Personnel Psychology, 56(1), 125–154.
- Gaugler, B. B., Rosenthal, D. B., Thornton, G. C., III, & Bentson, C. (1987). Meta-analysis of assessment center validity. Journal of Applied Psychology, 72(3), 493–511.
- Klimoski, R., & Brickner, M. (1987). Why do assessment centers work? The puzzle of assessment center validity. Personnel Psychology, 40(2), 243–260.
- Kuncel, N. R., & Sackett, P. R. (2014). Resolving the assessment center construct validity problem (as we know it). Journal of Applied Psychology, 99(1), 38–47.
- Lance, C. E. (2008). Why assessment centers do not work the way they are supposed to. Industrial and Organizational Psychology: Perspectives on Science and Practice, 1(1), 84–97.
- Lievens, F., & Thornton, G. C., III. (2005). Assessment centers: Recent developments in practice and research. In A. Evers, O. Smit-Voskuijl, & N. Anderson (Eds.), Handbook of personnel selection (pp. 243–264). Blackwell.
- Melchers, K. G., & König, C. J. (2008). It is not yet time to dismiss dimensions in assessment centers. Industrial and Organizational Psychology: Perspectives on Science and Practice, 1(1), 125–127.
- Sackett, P. R., & Dreher, G. F. (1982). Constructs and assessment center dimensions: Some troubling empirical findings. Journal of Applied Psychology, 67(4), 401–410.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- Thornton, G. C., III, & Rupp, D. E. (2006). Assessment centers in human resource management: Strategies for prediction, diagnosis, and development. Lawrence Erlbaum.
- Wirz, A., Melchers, K. G., Kleinmann, M., Lievens, F., Annen, H., Blum, U., & Ingold, P. V. (2020). Do overall dimension ratings from assessment centres show external construct-related validity? European Journal of Work and Organizational Psychology, 29(3), 405–420.