Most manager scorecards produce one number per manager. That's enough to rank them. It rarely tells anyone what to do on Monday. Below are 14 manager effectiveness metrics grouped into outcome, behavior and perception families, the formula for each, an honest account of which ones a manager controls, and the attribution problem that makes manager league tables so easy to misread.
Swipe to see the full team →
The 3.8 says nothing about who. The profiles say who ranks the two drivers that moved, and those two people need opposite things from the same manager.
Manager effectiveness metrics are the measures used to judge how well a manager leads a team. They come in three families: the outcomes the team produced, the management behaviors the manager performed, and how the team perceives being managed by them. A usable set spans all three, because each family on its own can be gamed, delayed, or explained away by circumstances the manager never controlled.
What the team produced: retention, promotions, goal attainment, ramp time. The metrics executives trust most. They are also the ones most contaminated by things the manager never chose.
What the manager did: one-to-ones held, feedback given, development conversations covered, decisions returned. The only family a manager can change this week by deciding to.
What it feels like to report to them: upward feedback, engagement, psychological safety, and whether each person's drivers are being met. The closest thing to a leading indicator.
Track five to seven manager effectiveness metrics: two outcome, two behavior, and two or three perception. Read the outcome metrics as direction, the behavior metrics as accountability, and the perception metrics as your early warning. Full breakdown of all 14 is below.
The population being measured is itself under strain. Gallup's 2026 State of the Global Workplace puts the scale of it plainly, and it is worth knowing before you design a scorecard that assumes the manager is the stable part of the system.
of managers were engaged at work in 2025, down five points from 27% the year before, the steepest single-year fall in the manager series Gallup publishes.
total decline in manager engagement since 2022. A manager scorecard built in 2022 is measuring a different population than the one you have now.
of managers were engaged inside best-practice organizations, close to quadruple the 20% recorded for employees globally. Our reading of that gap: where a manager works predicts more than who they are.
of the variance in team engagement is attributable to the manager, which is exactly why the measurement is worth getting right.
Low engagement cost the world economy roughly $10 trillion in lost productivity in 2025, about 9% of global GDP, on Gallup's estimate. Manager measurement is one of the few levers that sits upstream of that number.
Asked for a manager effectiveness dashboard by the end of the quarter
"I can build the scorecard. I just can't defend it yet in the room where it gets used."
Runs 40 to 400 people through a layer of managers
"I have twelve scores and no idea what to say differently to any of the twelve."
Every metric below includes how to calculate it and the conditions under which it reads wrong. That second part matters more than the formula: most manager scorecards fail because a metric was trusted in a situation it was never able to describe.
Read these as direction over several quarters. Any one of them, over one quarter, on a team of six, is noise.
The share of the manager's team who chose to leave over a period. The metric almost every scorecard opens with, and the one most often used unfairly.
Formula: (voluntary leavers from the team ÷ average team headcount) × 100, annualised Reads wrong when: the team is small, the market for those skills is hot, or the manager inherited a team already halfway out the door. On a team of six, one resignation is 17%.Of the people who left, how many you wanted to keep. This separates a manager losing strong people from a manager finally addressing a performance problem.
Formula: (regrettable leavers ÷ total voluntary leavers) × 100 Reads wrong when: the regrettable flag is set after the fact by the same manager being measured. Fix the incentive by having the flag set at exit by a second person.How often people move up or across out of this manager's team. Managers who grow people are visible here and often invisible everywhere else.
Formula: (promotions + lateral moves out of the team ÷ average team headcount) × 100, annualised Reads wrong when: the org has no open roles. A manager can develop people impeccably for two years and score zero because nothing above them moved.The proportion of the team's committed goals that were met in the period. The closest outcome metric to the business result the manager was hired to produce.
Formula: (goals met ÷ goals committed) × 100 Reads wrong when: the manager sets their own goals. Attainment then measures ambition calibration, and the reliable way to score 100% is to commit to less.How long a new joiner on this team takes to reach an agreed proficiency bar. A clean read on the manager's onboarding, because the clock starts on their watch.
Formula: median days from start date to the agreed proficiency milestone Reads wrong when: the proficiency bar is defined differently team to team, which it usually is. Compare a manager against their own trend rather than against another team.The only family a manager can move this week by deciding to. Use these for accountability. Keep the targets low enough that nobody has to game them.
Whether the one-to-ones on the calendar actually happen. Cheap to collect. When a manager goes underwater, this is the first thing they drop. We know of no study showing it predicts turnover on its own, and we have never seen a healthy team where these quietly stopped.
Formula: (one-to-ones held ÷ one-to-ones scheduled) × 100, per report per quarter Reads wrong when: you measure only the aggregate. A manager at 90% who has cancelled on the same person six times in a row looks fine. Track it per report.How often this manager gives specific, documented feedback, praise and correction both. Distinguishes managers who feed back continuously from managers who save it for review season.
Formula: documented feedback instances ÷ reports ÷ quarter Reads wrong when: the target becomes a quota. Volume rises, specificity collapses, and you have taught the team to ignore feedback. Pair it with a quality item in the survey.The share of the team who have had a real conversation about where they are going in the last six months. Coverage catches the person quietly skipped over.
Formula: (reports with a documented development conversation in the last 6 months ÷ total reports) × 100 Reads wrong when: the documentation is the deliverable. A filed form proves a form was filed. Ask the report whether the conversation changed anything.What proportion of the team was recognised at least once in the quarter. The raw count hides that answer, and reach exposes the manager with one favourite.
Formula: (reports recognised at least once in the quarter ÷ total reports) × 100 Reads wrong when: you track volume instead of reach. Ten recognitions all pointed at the same two people is a distribution problem wearing a good number.How long the team waits on this manager for a decision or an approval. Rarely put on a scorecard. In our experience it's the thing reports complain about most.
Formula: median elapsed time from a request reaching the manager to a decision being returned Reads wrong when: the manager lacks the authority to decide. You then measure the layer above them, which is worth knowing and unfair to attribute here.The nearest thing to an early warning, and the family where anonymity and team size need real care.
The composite of manager-specific items in an upward feedback survey: clarity, support, fairness, growth, and whether the person would choose to work for them again.
Formula: mean percentage favorable across the manager-specific items, reported only at n ≥ 4 Reads wrong when: reported below four respondents, or when a manager sits in the room while the team completes it. Both destroy the candor the index depends on.The team's own engagement result, which is where the 70% manager variance figure actually shows up. Useful, lagging, and averaged twice over.
Formula: percentage favorable, or mean on a 5-point scale, across the engagement items for that team Reads wrong when: read as a manager verdict. A reorganisation, a pay freeze or a cancelled product will move a team's engagement with no help from the manager. See how to measure employee engagement for the full method.Whether people on this team will admit a mistake, disagree with the manager, or raise a problem early. The upstream condition for every other metric on this page.
Formula: mean favorable across the psychological safety items, tracked per team over time Reads wrong when: the score is high because nobody has been tested yet. A team that has never had bad news to deliver has an untested safety score.The fit between the drivers each report ranks highest and what their role currently delivers. Two halves: motivator profiles are individual by design, so a manager knows who ranks Autonomy or Feedback top, while the motivation satisfaction pulse reports for the team at a minimum of three respondents, which keeps individual answers unattributable. Read together, a team-level dip points at named people.
Formula: each report's ranked drivers (individual) read against the team's satisfaction pulse on those drivers (team level, n ≥ 3) Reads wrong when: the pulse is read as an individual score. The pulse is a team reading whose own floor is three respondents; the four-report floor we recommend for the survey-based perception metrics is deliberately stricter. The profile is the individual half, and it is what tells you which people a team-level dip is most likely about.A list of 14 is a set of choices, and the choices are arguable. Here are the ones we expect to be challenged on.
Metric 11 is upward feedback: the manager's own reports rating the experience of being managed by them. A 360 widens that to peers, internal customers and the manager's own manager, which answers a different question and costs considerably more to run. For a scorecard you refresh quarterly, upward feedback is the right instrument, because the people best placed to judge day-to-day management are the people being managed. Keep the 360 for development cycles and promotion decisions, where the breadth earns its cost and the once-a-year cadence is acceptable.
The practical reason to keep the three families separate is that they answer different questions, arrive at different speeds, and are confounded to very different degrees. Mixing them into one composite score destroys all three properties at once.
| Family | Question it answers | How fast it moves | How confounded | Best used for |
|---|---|---|---|---|
| Outcome | Did this team produce the result? | Slow. Two to four quarters before a trend is real. | Heavily. Market, territory, headcount, inherited team, budget. | Direction over years. It cannot support a quarterly ranking. |
| Behavior | Did this manager do the job of managing? | Immediately. Visible within a fortnight. | Lightly. Mostly within the manager's control. | Accountability and coaching conversations. |
| Perception | What is it like to be managed by them? | Weeks, at survey cadence. Continuous if instrumented. | Moderately. Org events move it independently of the manager. | Early warning, and choosing who to talk to first. |
Scroll the table sideways to see every column →
Hold managers accountable for the behavior family, coach them using the perception family, and evaluate the organization using the outcome family. The common failure is doing it in reverse: holding a manager accountable for an outcome they did not control, while ignoring the behaviors they did.
Almost every manager effectiveness metric measures the manager plus their circumstances, and the two arrive fused together. A manager who took over a team mid-restructure, in a market where their skills are being bid up, on a product that just lost its funding, will produce worse numbers than a manager of equal skill who inherited none of that. Four things reduce the damage.
A league table across managers compares people with different starting conditions. A manager's own trend across four quarters holds most of those conditions constant, and it answers the question you care about: are they getting better?
The behavior family is the part the manager controls, so it carries the accountability. Outcome metrics belong on the scorecard as context and direction. Reversing that weighting produces a scorecard that rewards luck and inherited advantage.
Below about four reports, every rate metric becomes a coin flip and every survey result becomes identifiable. There isn't a clever way around this one. Report those teams qualitatively or roll them up. A precise-looking percentage on a team of three is worse than no number at all.
Log tenure in role, team composition at handover, and any reorganisation inside the window. Without those three fields, a low first-year score is indistinguishable from a low performer, and you will eventually act on the wrong one.
Not one of them was average, and no manager is either. In 1950 the US Air Force ran an anthropometric survey of its flying personnel, which left it with 131 available body measurements and a sample of 4,063 men. A lieutenant in the Aero Medical Laboratory, Gilbert Daniels, took ten of those measurements and put a plain question to the data: how many of these 4,063 men are average on all ten at once?
He was generous about what counted. His "approximately average" band was roughly the middle 30% of the population on each measurement, which he noted was "a considerably more generous portion of the group than is included by the exact average value." Then he went through the file one measurement at a time. Of the 4,063 men, 1,055 had approximately average stature. Of those 1,055, some 302 also had an average chest circumference. Of those 302, 143 also had an average sleeve length. Then 73. Then 28, 12, 6, 3, 2. At the tenth measurement, the number left was zero.
Daniels had picked the ten measurements that mattered for designing clothing, and he pointed out in passing that measurements for "cockpit layout or seat design could equally well have been chosen" and would have given much the same answer. He opened the note by calling the tendency to think in terms of the average man "a pitfall into which many persons blunder when attempting to apply human body size data to design problems." Earlier in the same note he calls the average man "a very rare specimen and very hard to fit." His conclusion is blunter: the average man is "a misleading and illusory concept as a basis for design criteria, and is particularly so when more than one dimension is being considered." The average man had been carefully measured, and he did not exist.
A manager effectiveness score of 3.8 is a garment cut for that man. It is a real measurement, honestly taken, describing somebody on that team who is not there. The two people about to resign are both inside the 3.8, and so is the person who is perfectly content, and the number gives the manager the same advice about all three.
Daniels published this in December 1952 as a short technical note titled "The 'Average Man'?", which is available in full as the DTIC copy on the Internet Archive. The elimination table is on page three, and the whole thing takes about ten minutes.
Cut to fit: one person's drivers, ranked and individually tracked
Five to seven. Two outcome, two behavior, and two or three perception. That is enough to triangulate and few enough that every manager can hold the whole set in their head. That is the real constraint.
Three metrics cannot triangulate. With one metric per family you have no way to tell a genuine signal from a quirk of that particular measure, and any single number can be moved without the underlying reality changing.
Past about eight, managers stop reading the dashboard and start asking which two you actually care about. You then have a fifteen-metric scorecard and a two-metric reality, and nobody has written down which two.
Name one metric as the anchor and say so out loud. Regrettable turnover on the manager's team is the usual choice. An unnamed anchor gets chosen for you, informally, by whoever runs the review meeting.
A metric that has not changed a single decision in a year is decoration. Drop it. The set should shrink and sharpen over time, and the annual review is where you check that a metric earned its slot.
This is the sequence that survives contact with a leadership meeting. It also answers the two questions people usually arrive with: how to evaluate manager success, and how to measure success as a manager when you are the manager being measured.
Decide what the scorecard is for before choosing a single metric: coaching, promotion, compensation, or spotting risk. Build a scorecard for coaching, then quietly use it for compensation, and you lose the honesty in every survey that feeds it.
Two outcome, two behavior, two perception, taken from the 14 above. Write the formula next to each one, in the document, in full. Half the disputes about a manager scorecard turn out to be disputes about a denominator.
A defensible split is roughly 50% behavior, 30% perception, 20% outcome. That ordering tracks how much control the manager has over each family, and it is the part you will be asked to justify.
Report each metric as the manager's own four-quarter trend, with the org median shown faintly for context. Ranking managers against each other invites the attribution error the whole scorecard exists to avoid.
State the minimum team size for reporting, who can see what, and what happens to a manager with three reports. Publishing this before the first run is what makes the survey answers usable. Retrofitting it afterwards does not restore the candor you already spent.
For each metric, write the sentence a manager should say in their next one-to-one when it moves the wrong way. A scorecard with no attached action produces anxiety and no behavior change. Most manager dashboards end up there and stay.
Six metrics for a manager with a team of six. Five of the six are things they can move inside a quarter. The sixth tells you whether the other five are working.
| Metric | Family | How it is calculated | Weight | What you do when it moves |
|---|---|---|---|---|
| Regrettable turnover rate (anchor) | Outcome | Regrettable leavers ÷ total voluntary leavers × 100 | 10% | Read the last two exit conversations before drawing any conclusion about the manager. |
| Internal mobility rate | Outcome | (Promotions + lateral moves out) ÷ avg headcount × 100 | 10% | Check whether roles existed to move into before crediting or blaming the manager. |
| One-to-one cadence adherence, per report | Behavior | One-to-ones held ÷ one-to-ones scheduled × 100 | 25% | Name the specific report who keeps getting cancelled on, and ask why that one. |
| Development conversation coverage | Behavior | Reports with a documented conversation in 6 months ÷ reports × 100 | 25% | Book the missing conversations this month, starting with the longest-tenured gap. |
| Manager effectiveness index | Perception | Mean % favorable across manager-specific survey items, n ≥ 4 | 15% | Find the individual item that moved and discuss that one. Leave the composite alone. |
| Motivator alignment, per report | Perception | Team satisfaction pulse (n ≥ 3) read against each report's ranked drivers | 15% | Look up who ranks the dipping driver highest, and talk to them this week. |
Scroll the table sideways to see every column →
Where this one does not apply. It assumes calendars, one-to-ones already in place, and a team large enough to report on. It breaks on frontline and shift work, where there is often no shared calendar and no company email; on spans of control past about twelve, where per-report tracking stops being realistic; and on teams of three or fewer, where the suppression floor removes most of the perception family. Unionised workplaces will usually need the behavior metrics agreed before they are collected at all.
Evaluate a manager against their own four-quarter trend. A peer ranking punishes whoever inherited the hardest team, so keep the comparison internal to the manager. Weight the behaviors they control above the outcomes they partly inherited. Record tenure in role, the team they took over, and any reorganisation inside the window, so a hard first year stays distinguishable from a weak manager. Suppress every number below four reports.
If you are the manager being measured, hold every one-to-one you scheduled, give each report documented feedback at least monthly, have a real development conversation with everyone inside six months, and return decisions inside a turnaround you have stated out loud. Those four are yours. Then watch your upward feedback score and whether each person's top drivers are being met, because those two move first when something is going wrong.
Behavior metrics monthly, since they come from calendar and workflow data you already hold and reading them costs nothing. Perception metrics quarterly at the slowest, and continuously if your instrument allows it. Outcome metrics annually, and only ever as a four-quarter trend. Reviewing an outcome metric monthly produces noise that people then act on, and that is worse than not reading it.
There is no universal target. A manager effectiveness metric is healthy when that manager's own four-quarter trend is flat or improving, and it is worth a conversation when it moves outside their own historical range. The three numbers we do stand behind are the suppression floor of four reports, a weighting of roughly 50% behavior / 30% perception / 20% outcome, and the arithmetic that one resignation on a team of six reads as 17%.
We publish no benchmark targets, for a reason worth defending. Tell an organization that 90% one-to-one adherence is good and you'll get 90% adherence with the same three people cancelled every month. That is the whole problem with a published target: it becomes a quota. Every target on this page would also be wrong for somebody, because the right feedback frequency for a team of three senior engineers is not the right frequency for twelve first-year hires.
Three comparisons carry more information than a benchmark, and you already have all three. Compare the manager to their own previous four quarters, which holds their circumstances roughly constant. Compare them to the median of managers inside the same function, which controls for the market they hire in. And compare each report against their own history rather than against their teammates, since a person whose engagement fell 20 points to a still-respectable score is the one worth talking to today.
People skills are usually declared unmeasurable, then measured badly by a single survey question about communication. They break into four observable components, and each one has a metric already on this page attached to it.
Adaptability is the component we see left off scorecards most often, largely because measuring it at all needs per-person data. Our deeper treatment of the skills themselves, rather than their measurement, is at how managers can improve their people skills. If you are measuring managers in order to develop them, that page is the other half of this one.
A team engagement score is already a mean across people. Rolling it into a manager composite averages it again. Two rounds of averaging removes the individual variation that told you what to do, and leaves a figure precise enough to act on and too coarse to act with.
Turnover, exit interviews and annual survey results all describe decisions people have already made. A scorecard built entirely from them is an accurate history of a problem, delivered after the window for fixing it closed.
A manager told their effectiveness index fell four points has learned that something's wrong and nothing about where. The useful version of that alert names the driver that stopped being met and the two people who rank it highest.
Every manager below the median gets sent the same coaching module. The manager who over-manages an autonomy-driven engineer and the manager who under-communicates with someone who needs feedback both receive "improve your communication", which helps neither.
Attuned adds one metric to the set on this page and changes the resolution of two others. Two instruments: a 10-minute assessment that ranks one person's 11 drivers, and a recurring pulse on whether those drivers are being met, read at team level with a floor of three respondents.
Each team member completes the Intrinsic Motivation Assessment: 55 forced-choice questions, about 10 minutes, producing a ranked profile across all 11 motivators. There are 39,916,800 ways to order 11 things, so across the 10,000+ people we have assessed, two identical profiles would be close to a coincidence. Your team has six of them, and the scorecard reports their mean.
The assessment: forced-choice, 55 questions
A manager score of 3.8 says nothing about who. When the team's satisfaction pulse on Autonomy falls, the manager already knows which two people rank Autonomy highest, because the profiles are individual. The dip arrives with a shortlist.
Attuned runs the satisfaction pulse on a short recurring cycle, so a decline surfaces in weeks instead of at the end of the year. That moves the perception family from confirming what already happened toward flagging what is about to happen.
AI TalkCoach reads each person's motivator profile and current gaps and gives the manager what to ask, how to frame it and what to steer around. It is the "attach one action to every metric" step from the build sequence, with the language already drafted.
Where a whole team's drivers cluster, the team view shows it, and the individual readings stay intact underneath. You get the pattern and keep the people. The composite score threw that away.
Satisfaction is tracked against the drivers a team actually ranks highest. A downward slope on one of them moves weeks before a resignation decision, which puts it ahead of the annual survey cycle, and it tells you the driver and roughly when the slide started.
Motivator satisfaction, tracked over time
A 20-minute call, your own manager scorecard on the screen, and an honest read on which metrics are carrying weight they cannot support.
No preparation needed. Bring the scorecard you already have.
Azon Recruitment Group is an Irish recruitment firm, so the metrics on their scorecard are not the ones a manufacturer or a bank would pick. What transfers is the shift underneath: from a manager's general impression of how someone was doing, to a specific per-person reading that could be checked. That shift is what the 14 metrics above are for.
An award-winning Irish recruitment agency, and one of the country's fastest-growing talent providers.
Azon's read on how people were doing rested on body language and on what surfaced through their managers, until a resignation proved the instinct wrong. "Typically, if an employee was engaged and they looked like they were happy, the assumption would be they were fine," says Denise Grant, Manager for HR Recruitment at Azon, "until you get their resignation and you realize at an exit interview that they weren't as happy as you had assumed." That is the perception family failing in the way it usually fails: the signal existed, and no instrument was pointed at it. Attuned gave them a measured, individual view of what each person valued, so a general impression could be checked against something specific.
"Being able to get to the nub of people's underlying motivations at the start of a process and see what really drives and motivates people in the workplace has been very helpful when trying to hire, and we've seen a dramatic increase in the numbers of people we hire that we feel we've gotten right, and that are a right fit for the business." Kevin Halligan, Associate Director, Banking & Financial Services, Azon
"Since beginning to use the software, we have really seen that benefit translating to earnings for our business." Kevin Halligan, Associate Director, Banking & Financial Services, Azon
A scorecard tells you which manager to talk to. These cover what to say once you get there.
Our own data set: motivator profiles from 10,000+ people across four generations, including how driver profiles differ by role and seniority. The empirical basis for metric 14.
Get the report →The qualities that show up behind a strong scorecard, and why the same behavior lands differently with two different reports.
Read: what makes a good manager →What a manager systematically fails to see about their own team, and the gap a perception metric exists to close.
Read: turning blindspots into strengths →The five failure modes behind a one-to-one cadence that looks compliant on the dashboard and produces nothing.
Read: why 1-on-1 meetings fail →How to measure metric 13 properly, including the items to use and why an untested high score is worth treating with suspicion.
Read: measuring psychological safety →The playbook behind the anchor metric: why regrettable turnover happens and which signals precede it.
Read: preventing unwanted turnover →Related pages: how to measure employee engagement for the team-level method, turnover metrics every CEO should track for the outcome family in depth, manager feedback tool for running the upward feedback that produces metric 11, and 1-on-1 coaching software for the conversation layer.
Every other metric on this page resolves to a team and stops there. Motivator profiles are individual by design, so when a team reading moves you can see which people rank that driver highest. The shortlist is the part a manager can act on.
Motivator satisfaction moves weeks before someone decides to leave, which puts it ahead of turnover, exit interviews and the annual survey. Those three report decisions already taken.
Attuned worked with psychologists to define 11 workplace motivators, each scored 0 to 100 against a single global norm, so a 72 means the same thing in Tokyo and in Texas. It sits in the same family of self-report instruments as the Reiss Motivation Profile (Reiss, 2004) and Amabile's Work Preference Inventory (1994). We have not published a peer-reviewed validation paper, and we do not claim one. The full instrument comparison is on our intrinsic motivation assessment page.
AI TalkCoach converts one person's profile and current gaps into prompts for the next one-to-one. The manager opens it while preparing, and the metric arrives with the sentence to say.
Manager effectiveness metrics are the measures used to judge how well a manager leads a team. They fall into three families: outcome metrics (team turnover, promotions, goal attainment, ramp time), behavior metrics (one-to-ones held, feedback given, development conversations covered, decision turnaround), and perception metrics (upward feedback, engagement, psychological safety, and how well each person's motivational drivers are met). A usable set draws from all three, because each family on its own can be gamed, delayed, or explained away by circumstances the manager never controlled.
Measure it on the things you control and read the rest as context. Concretely: hold every scheduled one-to-one, give documented feedback to each report at least monthly, have a real development conversation with every report inside six months, and return decisions inside a stated turnaround. Then watch two perception measures, your team's upward feedback score and whether each person's top drivers are being met, as your early warning. Team turnover and goal attainment still belong on the sheet. Just don't steer by them week to week: by the time either one moves, whatever moved it was decided months ago.
Evaluate a manager against their own trend across four quarters rather than against a ranking of their peers, using six metrics: two outcome, two behavior, two perception. Weight the behavior family heaviest, at roughly 50%, because that is the part the manager controls. Record what they inherited (tenure in role, team composition at handover, any reorganisation in the window) so a difficult first year is distinguishable from poor performance. Suppress reporting below four reports, where rate metrics become noise and survey answers become identifiable.
Five to seven: two outcome, two behavior, and two or three perception. Three cannot triangulate, since a single metric per family gives you no way to separate a real signal from a quirk of that measure. Past about eight, managers stop reading the dashboard and start asking which two you actually care about. Name one metric as the anchor, usually regrettable turnover on the manager's team, and review the whole set once a year, dropping any metric that has not changed a decision.
Break people skills into four observable components and attach an existing metric to each. Attention: one-to-one cadence adherence per report, plus whether the manager can name each report's top driver unprompted. Candor: feedback frequency paired with the psychological safety score, since high feedback with low safety means the delivery is causing harm. Adaptability: the team's motivator profiles read side by side, which exposes a manager running one style against six people who rank six different drivers highest. Follow-through: decision turnaround time and development conversation coverage, verified with the report rather than the file. Adaptability is the component we see left off scorecards most often, largely because measuring it at all needs per-person data.
A manager effectiveness scorecard is a short, fixed set of metrics reported per manager on a regular cadence, with the formula for each written down and a stated weighting between them. A defensible build states its purpose first (coaching, promotion, compensation or risk), takes two metrics from each family, weights behavior at about 50%, compares each manager to their own four-quarter trend, publishes the minimum team size for reporting, and attaches a specific action to every metric. Scorecards that skip the purpose statement tend to get built for coaching and then used for compensation, which removes the honesty from every survey feeding them.
Team turnover is a useful direction and a poor verdict. It measures the manager together with their circumstances: the market for those skills, the team they inherited, headcount and budget decisions taken above them, and the fortunes of the product they support. On a team of six, one resignation reads as 17%. Use regrettable turnover rather than raw turnover, have the regrettable flag set at exit by someone other than the manager being measured, track it over four quarters against that manager's own history, and give it a minority weighting on the scorecard.
Employee engagement describes the state of the people on a team. Manager effectiveness describes the practice of the person leading it. They are tightly linked, since Gallup attributes around 70% of the variance in team engagement to the manager, and they need separate measurement: a team's engagement can fall because of a pay freeze or a reorganisation with no contribution from the manager, and a manager can be doing the job well inside a difficult quarter. Engagement belongs on a manager's scorecard as one perception metric among several. Treat it as the verdict and you will eventually discipline the wrong manager. Our full method for the team-level side is at how to measure employee engagement.
There is no universal target, and any published benchmark becomes a quota. A manager effectiveness metric is healthy when that manager's own four-quarter trend is flat or improving, and it is worth a conversation when it moves outside their own historical range. Compare a manager to their own previous four quarters, to the median of managers inside the same function, and compare each report against their own history rather than against their teammates. The three numbers we do stand behind are a suppression floor of four reports, a weighting of roughly 50% behavior, 30% perception and 20% outcome, and the arithmetic that one resignation on a team of six reads as 17%.
Yes, for the behavior and outcome families. One-to-one cadence adherence, feedback frequency, development conversation coverage, recognition reach and decision turnaround all come from calendar, HRIS and workflow data you already hold, and so do turnover, promotion rate, goal attainment and ramp time. The perception family does require asking people something, whether through an upward feedback instrument, a pulse, or a motivator assessment. You can build the whole thing from system data. It'll work, and it has one blind spot: it will show you a manager hitting every prescribed behavior while their team quietly checks out.
Bring the manager scorecard you have. We will show you where the averaging is hiding something, which metrics are carrying weight they cannot support, and what a per-person reading adds to the set.
Talk to Attuned
A working session on the manager metrics you already track.
Book a Call → Get the State of Motivation Report Or read how managers improve their people skills →"If an employee was engaged and they looked like they were happy, the assumption would be they were fine." Denise Grant, HR Recruitment, Azon
Prefer email? attuned.ai