AI Didn't Break Leadership Science. It Just Scaled the Cracks.
- Stacey Force

- 3 days ago
- 5 min read

Most AI-driven hiring and leadership tools inherit an old assumption: that fit for a job can be captured by one instrument, applied the same way regardless of what stage that job, or the team around it, is in. That assumption was already failing before AI arrived. Jobs are being disassembled into roles and tasks, teams have become the real unit of work, and the criteria for success shift as a team moves through stages. AI doesn't correct for any of this. It just runs the same flawed model faster and at greater scale.
Psychologists and HR teams are moving fast to bring AI into hiring and leadership development. Faster screening, faster scoring, faster synthesis of interview and assessment data. The instinct makes sense. The problem is what's underneath it.
Leadership Science Was Built for a Target That Doesn't Hold Still
Traditional selection science assumes a stable criterion: a defined job, a known set of outcomes, a normal distribution of performance you can measure someone against. Leadership and team performance rarely work that way. Outcomes are often lopsided rather than evenly distributed, a handful of teams driving disproportionate results while most cluster far below. The people available for any given role are a narrow slice of who exists, which distorts what "high performer" even looks like in the data. And the feedback organizations get is shaped by survivorship: you learn the most about the teams that made it, not the ones that didn't get the chance.
None of that is new. What's new is the confidence AI brings to a measurement problem that was already unstable before a model got involved.
The Unit Leadership Science Measures Is Disappearing
Most leadership assessment still treats the job as the fixed backdrop it's measuring someone's fit against, looking at personality to determine how well they'll perform in that role. But the backdrop isn't holding still. Jobs are being pulled apart into discrete roles and tasks, redistributed across people, teams, and increasingly AI systems itself.
What's left in its place is the team. Team composition, team stage, and team dynamics now carry more predictive weight than any individual's fixed profile. A framework built to score a person against a job title and static role is measuring something that barely exists anymore.
A recent review of the social skills research literature ran into this same problem from a different angle. Researchers analyzed 756 published definitions spanning six decades and found 26 named concepts collapsing into 15 core ideas once the redundant and conflated ones were stripped out. The 15 concepts split cleanly into three categories:
Antecedents: the stable, trait-level attributes a person carries into any situation. Extraversion, agreeableness, general social intelligence. These shift slowly, if at all.
Determinants: the malleable capacities sitting between trait and action: knowledge, developed skill, and motivation. These can be built and coached.
Behaviors: the observable, enacted actions that move a goal forward. Managing an impression, resolving a conflict, adjusting a message mid-conversation.
The finding that matters most for leadership science: most of the confusion in the field wasn't from having too little data. It came from collapsing these three categories into one score. A trait got treated like a behavior. A malleable skill got measured as if it were fixed. Once researchers separated what was stable from what was learned from how people behave, the field's decades of scattered findings started making sense as three distinct, coherent stories instead of one muddled one.
Leadership science is running into that same wall now, at scale, with AI attached.
What Modern Leadership Science Requires
If the job isn't a stable unit, the assessment can't be either. A newly formed team needs role clarity and vision above almost anything else. A team further along needs leaders who can communicate differently than they did at formation. A team moving into scale needs a different set of behaviors entirely. Using the same instrument across all three treats three different problems as one.
Some of what predicts team performance holds steady regardless of stage. Some of it only matters because of where the team sits right now. Traits are the baseline. Behavior is the pulse. The two aren't interchangeable, and collapsing them into a single score is how a team ends up handed a personality report when what they needed was a stage-appropriate read on what's happening now.
Traits in context, manifesting as behavior. That distinction is the entire fix.
Norms, Not One Dataset
This is also why a single comparative benchmark across every team and every stage doesn't hold up. The thing being measured changes shape as the team moves. What can be built instead is a set of norms calibrated to stage and context, which is a much more honest form of rigor than one universal score applied everywhere.
The Tools Don't Need Reinventing. The Timing Does.
None of this is an argument against personality science or structured behavioral assessment. Those tools work. The failure isn't in the instrument, it's in applying the same one regardless of where a team is at a given moment.
Where AI Fits Into Leadership Science
Applied to a flexible, stage-aware approach, AI is useful. It can surface evidence faster, flag patterns across a larger set of teams, and help match the right tool to the right moment. What it shouldn't do is make the call. If AI is fed a model that already confuses trait with behavior, AI won't catch the mismatch. It will score faster against the wrong thing. Clean categories have to come before speed.
Leadership science doesn't need a faster version of the old approach. It needs one built around the team as the real unit of work, flexible enough to know which tool applies, and when. The unit of measurement is being rebuilt right now. That's not a problem to patch. It's a design opportunity, and whoever builds the norms first sets the standard everyone else measures against.
FAQs
Why can't AI just improve the accuracy of existing leadership assessments? Accuracy isn't the core problem. The issue is that most assessments read someone's fit against a static job or role, and that backdrop is dissolving as work gets broken into tasks and redistributed across teams. AI can process data faster, but it can't correct for reading fit against the wrong backdrop to begin with.
What should organizations use instead of a single leadership assessment? A stage-aware approach that matches the instrument to what the team needs right now, whether that's role clarity in a newly formed team, communication shifts in a maturing one, or different behaviors entirely as a team scales.
Is this an argument against personality assessments or behavioral science? No. The underlying tools hold up. The failure is applying one tool the same way at every stage instead of matching the right measure to the right moment.
How does team stage affect what predicts performance? Some characteristics predict performance regardless of stage. Others only matter because of where a team currently sits. Treating those as interchangeable is how organizations end up measuring the wrong thing at the wrong time.
.png)



Comments