When the Metric Becomes the Target

What happens to a design team the moment a number turns into a goal

9 min read

Give a team a number to hit and they will hit it. That is the good news and the whole problem. A metric is a shadow of something you care about, cast at a useful angle. Turn the shadow into the target and people learn to move the shadow without moving the thing. I have watched a design team raise engagement by 30 percent and make the product worse, and every dashboard said we were winning.

A spring scale hanging from its hook, weighing a load by how far an internal spring stretches.
A spring scale does not measure weight at all. It measures how far a spring stretched, and we agree to read that as weight. It works right up until someone leans on the pan, and that gap between the signal and the thing you actually care about is exactly where Goodhart's law lives.Bestalex, via Wikimedia CommonsCC0

Picture a hospital that gets judged on how long people wait in Accident and Emergency. A four hour limit sounds humane. So the ambulances start queuing outside the doors, patients held in the bay, because the clock only starts once you are through them. The wait did not shrink. It moved to a place the metric could not see. The number improved. The care did not. That is the pattern, and it does not care whether you run a hospital or a checkout flow.

Goodhart's law, and who actually said the famous line

Charles Goodhart was a monetary economist. In 1975 he made a dry observation about central banks. Any measure the bank tried to control, he noticed, would stop behaving the moment it became the thing under control. His own words were careful and a little clumsy: any observed statistical regularity tends to collapse once pressure is placed upon it for control purposes. That is the original. It is worth reading twice, because almost nobody quotes it.

The line everyone quotes came later, from an anthropologist. In a 1997 paper about how British universities were being audited, Marilyn Strathern compressed Goodhart into one sentence. People credit it to him. She wrote it. Getting this right matters, because the crisp version is the one that travels, and the person who made it travel was studying exactly our problem: what happens to human beings when their work is turned into a score.

When a measure becomes a target, it ceases to be a good measure.

Marilyn StrathernImproving ratings: audit in the British University system, 1997

Read it as a warning about incentives, not about maths. The measure does not rot on its own. It rots because you started rewarding it. The instant a person's raise, or a squad's roadmap, depends on a number, that number stops measuring the world and starts measuring the person's effort to move the number. Those are different things. They agree right up until the day someone finds the cheaper path.

A signal is not a target

The fix starts with a distinction most teams never draw. A signal is a reading you watch to understand what is happening. A target is a number you have promised to move. The same metric can be either. Time on page is a fine signal. As a target it is poison, because you can raise it two ways: make the content compelling, or make the interface so confusing that people cannot find the exit. Both push the graph up. Only one is worth having.

A metric used as a signal
  • You watch it to learn what changed
  • It sits next to three others that would catch a lie
  • A rise triggers a question, not a celebration
  • Nobody's bonus depends on it
  • You are allowed to conclude it went up for a bad reason
The same metric used as a target
  • You have promised leadership it will go up
  • It is reported alone, stripped of context
  • A rise ends the conversation
  • A team is compensated on it
  • Explaining a rise away is treated as making excuses
Nothing about the metric changes between these columns. What changes is what you are allowed to conclude, and that is the whole difference between learning and theatre.

This is why engagement is the most dangerous number in product design. It is the one most likely to rise for the wrong reason. A user who is delighted and a user who is lost both generate sessions, scrolls and clicks. The dashboard cannot tell them apart. Optimise engagement hard enough and you will build a product that is sticky the way a swamp is sticky. People are in it a long time and none of them are happy.

Improve the product

Exploit the metric

Pick a metric
e.g. time on page

Attach it to
a team's goals

People optimise
the metric directly

Easiest way
to move it?

Metric and reality
rise together

Metric rises,
reality falls

Dashboard says
you are winning

Nobody looks
closer

The failure is not the metric. It is the arrow from the metric to a team's incentives, which quietly rewards whichever path up is cheapest.

Google's research team built a way out of this that I lean on constantly. It is called HEART, and its real contribution is not the five categories. It is the discipline of going from goal to signal to metric, in that order, and never skipping to the metric. You state the goal in plain words. You decide what observable behaviour would signal progress. Only then do you pick a number. Do it backwards, starting from whatever is easy to count, and you get a target with no goal behind it.

The HEART framework from Rodden, Hutchinson and Fu at Google. The columns read left to right for a reason. Start at the metric and you have skipped the only two steps that make it mean anything.
CategoryThe goal, in wordsA signal you might watchA metric
HappinessPeople feel the tool respects themSurvey sentiment after a taskCSAT on task completion
EngagementPeople use the depth, not just the doorActions per active sessionMedian features touched per week
AdoptionNew people reach first value fastReached the core action oncePercent activated in seven days
RetentionPeople come back because it helpedReturns in a later weekWeek four retention rate
Task successPeople finish what they came to doCompleted the flow unaidedCompletion rate and time on task
The HEART framework from Rodden, Hutchinson and Fu at Google. The columns read left to right for a reason. Start at the metric and you have skipped the only two steps that make it mean anything.Google Research, CHI 2010

The average is a liar you trust

Here is the mistake I made for years. I reported means. Mean task time, mean latency, mean satisfaction. A mean is a single number pretending to speak for everyone, and it speaks loudest for the people in the middle, who were never your problem. The people you are failing live in the tail, and the tail is exactly what an average erases.

Think of a bus that is on time on average. Half the days it is five minutes early, half the days it is thirty five minutes late. The average is respectable. The experience is misery, because nobody rides the average bus. They ride the late one and miss the meeting. A mean task time of nine seconds can hide a group for whom the task takes two minutes and often fails. That group churns, quietly, and your headline number never flinches.

Median (p50) task time9 s
75th percentile (p75)14 s
95th percentile (p95)41 s
99th percentile (p99)88 s
A worked example of one flow. The mean here lands near 13 seconds and looks healthy. The story is at p95 and p99, where a real slice of users is stuck for over half a minute. Percentiles are the honest view.Worked example, Kousik Dutta

This is why Google's own site reliability practice reports percentiles, not averages, and why Core Web Vitals judge a site at the 75th percentile of its real users, not the mean. The rule I now follow: an average is a starting question, never an answer. If someone shows me a mean without a p95 next to it, I assume the p95 is bad and they have not looked. Usually I am right.

The users you never measured

There is a subtler lie underneath the average. Your analytics only contain the people who stayed long enough to be counted. The ones who bounced in the first five seconds, hit a wall, and left are barely in your data at all. So you tune the product for the survivors and call it listening to users. It is survivorship bias, and product analytics is riddled with it.

The classic version comes from the second world war. Analysts looked at bombers that came back full of holes and wanted to armour the parts with the most holes. The statistician Abraham Wald pointed out the obvious thing everyone had missed. The planes in front of them were the ones that survived. The holes they were not seeing, in the engines, were the holes that brought planes down. Armour where the data is empty. Your churn is the plane that did not come back, and it leaves no holes to count.

~85%
of usability problems surfaced by testing with just five users, in the classic finding
Nielsen Norman Group
75th
percentile is where Core Web Vitals grades a real site, precisely to stop the average hiding a bad tail
web.dev
Five people watched closely will show you what a million anonymous events cannot: the reason someone gave up.

This is the qualitative counterweight, and it is not soft. Five usability sessions catch the wall that made someone quit, because you watch them hit it. A million events will never tell you why, only that a line went down. Jakob Nielsen has since been clear that the five user figure is a guide for a single round of testing, not a magic constant, and that you test again after each fix. The point stands. The number tells you that something is wrong. A person tells you what.

Data driven, or data justified

Now the argument I will pick with the industry. Most of what gets called data driven design is data justified design. The two look identical in a deck and are opposites in practice. Data driven means the number could have changed your mind, and sometimes did. Data justified means the decision was made, and the number was chosen afterwards to defend it. One is science. The other is a lawyer building a case.

There is a clean test to tell them apart. Before you look at the result, write down what would make you kill the idea. Name the number and the threshold. If you cannot, you are not measuring, you are searching for a quote that agrees with you. A decision metric is one that could talk you out of shipping. A vanity metric is one that only ever goes up and only ever confirms you. Total registered users is vanity. It cannot fall and it cannot change a plan. Week four retention is a decision metric. It can ruin your afternoon, which is exactly why it is worth watching.

You can hear the difference in a review. Data justified design sounds like this: we knew the redesign was right, and look, sign ups went up two points. Nobody asks what else changed that month, or whether a holiday sale did the work. Data driven design sounds duller and more honest: we said we would only ship if activation held above 40 percent, it came in at 43, so we shipped. The first is a story with a number stapled on. The second names the threshold before the result, which is the only order that keeps you honest.

So what do you do on Monday. Pick your one or two real goals and write them as sentences a user would recognise. For each, choose a signal, then a metric, in that order. Report every headline number with a percentile beside it. Keep at least one guardrail that is allowed to say no. And before any test, write the result that would change your mind, and keep the note. If you never write that note, you already know the answer you are going to find.

Was this useful? Your choice stays private to this device.