Insights

Leadership Assessment Results: Why Low Scores Aren’t Your Action Plan

A low score in a leadership assessment or 360 is a signal, not a cause. Read results as Score → Signal → Hypothesis → Context → Action.
Leadership Assessment Results: Why Low Scores Aren’t Your Action Plan
Telta team
2026-10-11
Telta team
|
2026-10-11
Contents

The riskiest moment in a leadership assessment may not be when a score comes back low. It may be the moment you think you already know why, and what to do about it.

You find the lowest scores in a leadership assessment:

Communication: 3.0
Developing others: 2.9

The natural next question is, “So what do we fix?” But there is something to check first. You don’t yet know what the low score is a signal of.

The Short Answer
‍
A low score is not a cause. It is a signal. Interpreting results step by step, Score → Signal → Hypothesis → Context → Action, keeps you from chasing the wrong development goal.

‍

Leaders Who Failed Showed Common Signals in Their 360 Feedback

Jack Zenger and Joseph Folkman used 360-degree feedback data to study how leaders derail. In one study, they collected data on more than 450 Fortune 500 executives and examined the 31 who were fired within the next three years. Separately, they analyzed the bottom 10% of more than 11,000 leaders, those rated least effective.

450+

Fortune 500 executives with 360 data

31

Executives fired within three years

11K+

Leaders in a separate dataset

Compare common leadership flaws → look for weak signals that appear before failure

A summary of the study design described in HBR by Zenger & Folkman (2009).

The 360 reports did not say, “This executive will be fired in three years.” What they contained were several weak signals that could be linked to the failures that followed.

A leadership assessment plays a similar role. It is less about predicting the future and more about telling apart the scores you can safely move past from the signals that deserve a closer look.

Source: Harvard Business Review, “Ten Fatal Flaws That Derail Leaders”

‍

The Same 3.0 in Communication Can Mean Very Different Problems

The easiest reading is “We need to talk more.” So you add one-on-ones or team meetings. Push that logic a little further and you schedule more team dinners.

Observed scoreCommunication 3.0
Information flow

Important information is shared late.

Decision-making

Conclusions and the reasons behind them aren’t shared.

Voice and safety

It’s hard to disagree openly.

Listening

The leader talks a lot but doesn’t listen enough.

Feedback quality

Plenty of feedback, but the next step is unclear.

Alignment

Direction from above never reaches the team.

Different causes can sit beneath the same low communication score. Different causes call for different remedies.

A low communication score is an observed result, not an explanation of its cause.

‍

Wells Fargo: A Leadership Failure Whose Visible Problem Hid the Real Cause

In the Wells Fargo sales practices scandal, thousands of employees were fired for improper sales conduct, and millions of potentially unauthorized accounts came to light. Materials from the Federal Reserve and its Office of Inspector General (OIG) later pointed to aggressive sales goals, pressure from management, a failure to escalate risks properly, and governance problems.

The OIG noted in particular that senior management and the board tended to view the problem as individual employee misconduct rather than as a flaw in the sales model itself.

Surface reading“A few bad employees”

The cause is explained through the individuals involved.

What investigations foundSales goals + pressure + control failures

Incentives and management and risk structures amplified the problem.

Thought experimentEven if a leadership assessment had been in place, it could not have warned, “Unauthorized accounts are being opened.” It might, however, have surfaced weak signals around unrealistic targets, adherence to principles, speaking up, and passing bad news upward. This does not mean such assessment data actually existed.

A thought experiment linking actual investigation findings to leadership assessment. The Wells Fargo case does not prove that leadership assessments work, and no hypothetical results are presented as fact.

Major leadership failures don’t always arrive as a sudden crisis with no warning. There may have been small signals whose meaning was read too narrowly or whose importance was underestimated.

Source: Federal Reserve (2018) · Federal Reserve OIG (2020) · HBR (2016)

‍

Differences Between Raters May Not Be Noise

Ron Carucci wrote in HBR about an executive who received sharply conflicting feedback from about 25 stakeholders. Some described the executive as supportive and caring; others called them self-centered. Some felt trusted with autonomy; others felt micromanaged.

ASupportive

Helps when it matters.

BSelf-centered

Sticks to their own view.

CGrants autonomy

Doesn’t interfere in the details.

DMicromanages

Feels overly involved.

Before averaging it into “Empowerment: 3.5” → ask why experiences diverge

Conflicting feedback isn’t noise to be averaged away. It can be an important clue that people experience the leader differently in different situations.

Source: Harvard Business Review, “How to Make Sense of Conflicting Feedback on Your Leadership”

‍

What to Check First When Rater Scores Differ

Say a leader rates themselves 4.3 while their team rates them 3.1. Rather than concluding on the spot that the leader lacks self-awareness, it’s safer to check at least these four things.

Self vs. othersIs the leader the only one who sees it differently?

Look at which behavioral items drive the gap between self-ratings and others’ ratings.

Manager vs. teamDoes the experience depend on direction?

Behavior when managing up can differ from behavior when leading a team.

Specific groupsDoes everyone feel the same problem?

Check whether the gap appears only in certain teams, roles, or relationships.

Behavioral itemsWhere do ratings split at the behavior level?

Look past the overall communication score to observable behaviors like listening, sharing information, and explaining decisions.

An important limitation360-degree feedback shows how people experience and perceive a leader’s behavior. Scores alone cannot prove that a specific behavior directly caused a particular business outcome.

‍

How Big a Gap Counts as Meaningful?

There is no universal threshold, such as 0.3 or 0.5 points, that applies to every leadership assessment. The same gap can mean different things depending on the scale, item reliability, number of raters, norms, confidence intervals, and the distribution of results.

In practice, it helps to look at more than a single number:

  • Does the direction repeat? Do gaps point the same way across multiple items?
  • Is it concentrated in one group? Do direct reports consistently rate the leader lower than others do?
  • Can it be explained at the behavior level? Does a pattern show up in specific behavioral items, not just abstract competencies?
  • Does it fit the context? Do written comments or real work situations show similar signals?

A large gap doesn’t automatically mean a bigger problem, and a small gap isn’t automatically safe to ignore.

‍

Why You Need More Steps Between Score and Action

Five steps to read leadership assessment results: Score, Signal, Hypothesis, Context, Action
A five-step framework for interpreting leadership assessment results: Score → Signal → Hypothesis → Context → Action

Nor does 360-degree feedback turn into change on its own. In a meta-analysis of 24 longitudinal studies, Smither, London, and Reilly found that rating improvements over time were generally small, and that factors such as recognizing the need to change, accepting the feedback, setting goals, and taking action mattered. A 2026 HBR article likewise stressed the value of discussing 360 results with colleagues after receiving them.

Source: Smither, London & Reilly (2005) · HBR (2026)

A good leadership assessment doesn’t rush to tell you what to fix. It tells you which signals you shouldn’t ignore.

‍

Telta Focuses on Closing the Interpretation Gap, Not Handing Down Answers

If an AI spots a low leadership score and immediately concludes, “Take a communication course,” it repeats the very mistake this article describes.

In Telta’s leadership assessment reports and Report Assistant, what matters is making it easier to explore results within the report, check differences between rater groups and detailed results, and narrow down what to ask next.

Explore results

Quickly find the results you need within the report.

Check context

Review detailed results alongside rater differences.

Ask the next question

Sharpen the questions to verify instead of jumping to a cause.

We make no claims about product performance or accuracy. The focus is on supporting how reports are interpreted.

‍

FAQ: Interpreting Leadership Assessment Results

1. Should We Start With the Lowest Score?

Not necessarily. A low score is first a signal to investigate. It’s safer to review detailed items, differences between rater groups, written comments, and work context, test possible explanations, and then set priorities.

2. Why Do 360 Scores Differ So Much Between Raters?

Raters may interact with the leader in different situations and relationships. Before averaging scores, check which groups report which differences in experience. That often reveals important context.

3. How Do We Turn Leadership Assessment Results Into Action?

Rather than jumping from Score to Action, define the Signal, form possible Hypotheses, check rater-level results and real work Context, and then decide on Action.

4. How Should We Read Gaps Between Self-Ratings and Team Ratings?

Avoid concluding right away that the leader lacks self-awareness. First check which behavioral items show the gap, whether the same pattern appears with other rater groups, and what the sample size and response context look like.

5. Whose Ratings Matter Most: Managers, Peers, or Direct Reports?

No rater group is always more important. It depends on who actually has the chance to observe the behavior in question and on the purpose of the assessment.

6. Is a 0.3- or 0.5-Point Difference Meaningful?

There is no universal threshold that applies to every assessment. It depends on the scale, reliability, number of raters, norms, and confidence intervals, so look at whether the direction repeats and at detailed behavioral patterns.

7. Can AI Automatically Find the Cause of Leadership Assessment Results?

Pinning down the real cause from assessment data alone is risky. Use AI as a tool to explore and compare results and to narrow down what to verify next, and make the final judgment on causes with organizational context and further checks.

References

  1. Zenger & Folkman, HBR (2009)
  2. Federal Reserve, Wells Fargo enforcement action (2018)
  3. Federal Reserve OIG, Wells Fargo investigation summary (2020)
  4. Ron Carucci, HBR (2022)
  5. Smither, London & Reilly, Personnel Psychology (2005)
  6. Brenda Steinberg, HBR (2026)

Related Reading

‍

Talk to Us About Leadership 360