The riskiest moment in a leadership assessment may not be when a score comes back low. It may be the moment you think you already know why, and what to do about it.
You find the lowest scores in a leadership assessment:
The natural next question is, “So what do we fix?” But there is something to check first. You don’t yet know what the low score is a signal of.
The Short Answer
A low score is not a cause. It is a signal. Interpreting results step by step, Score → Signal → Hypothesis → Context → Action, keeps you from chasing the wrong development goal.
Jack Zenger and Joseph Folkman used 360-degree feedback data to study how leaders derail. In one study, they collected data on more than 450 Fortune 500 executives and examined the 31 who were fired within the next three years. Separately, they analyzed the bottom 10% of more than 11,000 leaders, those rated least effective.
Fortune 500 executives with 360 data
Executives fired within three years
Leaders in a separate dataset
A summary of the study design described in HBR by Zenger & Folkman (2009).
The 360 reports did not say, “This executive will be fired in three years.” What they contained were several weak signals that could be linked to the failures that followed.
A leadership assessment plays a similar role. It is less about predicting the future and more about telling apart the scores you can safely move past from the signals that deserve a closer look.
Source: Harvard Business Review, “Ten Fatal Flaws That Derail Leaders”
The easiest reading is “We need to talk more.” So you add one-on-ones or team meetings. Push that logic a little further and you schedule more team dinners.
Important information is shared late.
Conclusions and the reasons behind them aren’t shared.
It’s hard to disagree openly.
The leader talks a lot but doesn’t listen enough.
Plenty of feedback, but the next step is unclear.
Direction from above never reaches the team.
Different causes can sit beneath the same low communication score. Different causes call for different remedies.
A low communication score is an observed result, not an explanation of its cause.
In the Wells Fargo sales practices scandal, thousands of employees were fired for improper sales conduct, and millions of potentially unauthorized accounts came to light. Materials from the Federal Reserve and its Office of Inspector General (OIG) later pointed to aggressive sales goals, pressure from management, a failure to escalate risks properly, and governance problems.
The OIG noted in particular that senior management and the board tended to view the problem as individual employee misconduct rather than as a flaw in the sales model itself.
The cause is explained through the individuals involved.
Incentives and management and risk structures amplified the problem.
A thought experiment linking actual investigation findings to leadership assessment. The Wells Fargo case does not prove that leadership assessments work, and no hypothetical results are presented as fact.
Major leadership failures don’t always arrive as a sudden crisis with no warning. There may have been small signals whose meaning was read too narrowly or whose importance was underestimated.
Source: Federal Reserve (2018) · Federal Reserve OIG (2020) · HBR (2016)
Ron Carucci wrote in HBR about an executive who received sharply conflicting feedback from about 25 stakeholders. Some described the executive as supportive and caring; others called them self-centered. Some felt trusted with autonomy; others felt micromanaged.
Helps when it matters.
Sticks to their own view.
Doesn’t interfere in the details.
Feels overly involved.
Conflicting feedback isn’t noise to be averaged away. It can be an important clue that people experience the leader differently in different situations.
Source: Harvard Business Review, “How to Make Sense of Conflicting Feedback on Your Leadership”
Say a leader rates themselves 4.3 while their team rates them 3.1. Rather than concluding on the spot that the leader lacks self-awareness, it’s safer to check at least these four things.
Look at which behavioral items drive the gap between self-ratings and others’ ratings.
Behavior when managing up can differ from behavior when leading a team.
Check whether the gap appears only in certain teams, roles, or relationships.
Look past the overall communication score to observable behaviors like listening, sharing information, and explaining decisions.
There is no universal threshold, such as 0.3 or 0.5 points, that applies to every leadership assessment. The same gap can mean different things depending on the scale, item reliability, number of raters, norms, confidence intervals, and the distribution of results.
In practice, it helps to look at more than a single number:
A large gap doesn’t automatically mean a bigger problem, and a small gap isn’t automatically safe to ignore.

Nor does 360-degree feedback turn into change on its own. In a meta-analysis of 24 longitudinal studies, Smither, London, and Reilly found that rating improvements over time were generally small, and that factors such as recognizing the need to change, accepting the feedback, setting goals, and taking action mattered. A 2026 HBR article likewise stressed the value of discussing 360 results with colleagues after receiving them.
Source: Smither, London & Reilly (2005) · HBR (2026)
A good leadership assessment doesn’t rush to tell you what to fix. It tells you which signals you shouldn’t ignore.
If an AI spots a low leadership score and immediately concludes, “Take a communication course,” it repeats the very mistake this article describes.
In Telta’s leadership assessment reports and Report Assistant, what matters is making it easier to explore results within the report, check differences between rater groups and detailed results, and narrow down what to ask next.
Quickly find the results you need within the report.
Review detailed results alongside rater differences.
Sharpen the questions to verify instead of jumping to a cause.
We make no claims about product performance or accuracy. The focus is on supporting how reports are interpreted.
Not necessarily. A low score is first a signal to investigate. It’s safer to review detailed items, differences between rater groups, written comments, and work context, test possible explanations, and then set priorities.
Raters may interact with the leader in different situations and relationships. Before averaging scores, check which groups report which differences in experience. That often reveals important context.
Rather than jumping from Score to Action, define the Signal, form possible Hypotheses, check rater-level results and real work Context, and then decide on Action.
Avoid concluding right away that the leader lacks self-awareness. First check which behavioral items show the gap, whether the same pattern appears with other rater groups, and what the sample size and response context look like.
No rater group is always more important. It depends on who actually has the chance to observe the behavior in question and on the purpose of the assessment.
There is no universal threshold that applies to every assessment. It depends on the scale, reliability, number of raters, norms, and confidence intervals, so look at whether the direction repeats and at detailed behavioral patterns.
Pinning down the real cause from assessment data alone is risky. Use AI as a tool to explore and compare results and to narrow down what to verify next, and make the final judgment on causes with organizational context and further checks.
References
Related Reading