When the Metric Becomes the Target
What happens to a design team the moment a number turns into a goal
Give a team a number to hit and they will hit it. That is the good news and the whole problem. A metric is a shadow of something you care about, cast at a useful angle. Turn the shadow into the target and people learn to move the shadow without moving the thing. I have watched a design team raise engagement by 30 percent and make the product worse, and every dashboard said we were winning.

Picture a hospital that gets judged on how long people wait in Accident and Emergency. A four hour limit sounds humane. So the ambulances start queuing outside the doors, patients held in the bay, because the clock only starts once you are through them. The wait did not shrink. It moved to a place the metric could not see. The number improved. The care did not. That is the pattern, and it does not care whether you run a hospital or a checkout flow.
Goodhart's law, and who actually said the famous line
Charles Goodhart was a monetary economist. In 1975 he made a dry observation about central banks. Any measure the bank tried to control, he noticed, would stop behaving the moment it became the thing under control. His own words were careful and a little clumsy: any observed statistical regularity tends to collapse once pressure is placed upon it for control purposes. That is the original. It is worth reading twice, because almost nobody quotes it.
The line everyone quotes came later, from an anthropologist. In a 1997 paper about how British universities were being audited, Marilyn Strathern compressed Goodhart into one sentence. People credit it to him. She wrote it. Getting this right matters, because the crisp version is the one that travels, and the person who made it travel was studying exactly our problem: what happens to human beings when their work is turned into a score.
When a measure becomes a target, it ceases to be a good measure.
Read it as a warning about incentives, not about maths. The measure does not rot on its own. It rots because you started rewarding it. The instant a person's raise, or a squad's roadmap, depends on a number, that number stops measuring the world and starts measuring the person's effort to move the number. Those are different things. They agree right up until the day someone finds the cheaper path.
A signal is not a target
The fix starts with a distinction most teams never draw. A signal is a reading you watch to understand what is happening. A target is a number you have promised to move. The same metric can be either. Time on page is a fine signal. As a target it is poison, because you can raise it two ways: make the content compelling, or make the interface so confusing that people cannot find the exit. Both push the graph up. Only one is worth having.
- You watch it to learn what changed
- It sits next to three others that would catch a lie
- A rise triggers a question, not a celebration
- Nobody's bonus depends on it
- You are allowed to conclude it went up for a bad reason
- You have promised leadership it will go up
- It is reported alone, stripped of context
- A rise ends the conversation
- A team is compensated on it
- Explaining a rise away is treated as making excuses
This is why engagement is the most dangerous number in product design. It is the one most likely to rise for the wrong reason. A user who is delighted and a user who is lost both generate sessions, scrolls and clicks. The dashboard cannot tell them apart. Optimise engagement hard enough and you will build a product that is sticky the way a swamp is sticky. People are in it a long time and none of them are happy.
Google's research team built a way out of this that I lean on constantly. It is called HEART, and its real contribution is not the five categories. It is the discipline of going from goal to signal to metric, in that order, and never skipping to the metric. You state the goal in plain words. You decide what observable behaviour would signal progress. Only then do you pick a number. Do it backwards, starting from whatever is easy to count, and you get a target with no goal behind it.
| Category | The goal, in words | A signal you might watch | A metric |
|---|---|---|---|
| Happiness | People feel the tool respects them | Survey sentiment after a task | CSAT on task completion |
| Engagement | People use the depth, not just the door | Actions per active session | Median features touched per week |
| Adoption | New people reach first value fast | Reached the core action once | Percent activated in seven days |
| Retention | People come back because it helped | Returns in a later week | Week four retention rate |
| Task success | People finish what they came to do | Completed the flow unaided | Completion rate and time on task |
The average is a liar you trust
Here is the mistake I made for years. I reported means. Mean task time, mean latency, mean satisfaction. A mean is a single number pretending to speak for everyone, and it speaks loudest for the people in the middle, who were never your problem. The people you are failing live in the tail, and the tail is exactly what an average erases.
Think of a bus that is on time on average. Half the days it is five minutes early, half the days it is thirty five minutes late. The average is respectable. The experience is misery, because nobody rides the average bus. They ride the late one and miss the meeting. A mean task time of nine seconds can hide a group for whom the task takes two minutes and often fails. That group churns, quietly, and your headline number never flinches.
This is why Google's own site reliability practice reports percentiles, not averages, and why Core Web Vitals judge a site at the 75th percentile of its real users, not the mean. The rule I now follow: an average is a starting question, never an answer. If someone shows me a mean without a p95 next to it, I assume the p95 is bad and they have not looked. Usually I am right.
The users you never measured
There is a subtler lie underneath the average. Your analytics only contain the people who stayed long enough to be counted. The ones who bounced in the first five seconds, hit a wall, and left are barely in your data at all. So you tune the product for the survivors and call it listening to users. It is survivorship bias, and product analytics is riddled with it.
The classic version comes from the second world war. Analysts looked at bombers that came back full of holes and wanted to armour the parts with the most holes. The statistician Abraham Wald pointed out the obvious thing everyone had missed. The planes in front of them were the ones that survived. The holes they were not seeing, in the engines, were the holes that brought planes down. Armour where the data is empty. Your churn is the plane that did not come back, and it leaves no holes to count.
This is the qualitative counterweight, and it is not soft. Five usability sessions catch the wall that made someone quit, because you watch them hit it. A million events will never tell you why, only that a line went down. Jakob Nielsen has since been clear that the five user figure is a guide for a single round of testing, not a magic constant, and that you test again after each fix. The point stands. The number tells you that something is wrong. A person tells you what.
Data driven, or data justified
Now the argument I will pick with the industry. Most of what gets called data driven design is data justified design. The two look identical in a deck and are opposites in practice. Data driven means the number could have changed your mind, and sometimes did. Data justified means the decision was made, and the number was chosen afterwards to defend it. One is science. The other is a lawyer building a case.
There is a clean test to tell them apart. Before you look at the result, write down what would make you kill the idea. Name the number and the threshold. If you cannot, you are not measuring, you are searching for a quote that agrees with you. A decision metric is one that could talk you out of shipping. A vanity metric is one that only ever goes up and only ever confirms you. Total registered users is vanity. It cannot fall and it cannot change a plan. Week four retention is a decision metric. It can ruin your afternoon, which is exactly why it is worth watching.
You can hear the difference in a review. Data justified design sounds like this: we knew the redesign was right, and look, sign ups went up two points. Nobody asks what else changed that month, or whether a holiday sale did the work. Data driven design sounds duller and more honest: we said we would only ship if activation held above 40 percent, it came in at 43, so we shipped. The first is a story with a number stapled on. The second names the threshold before the result, which is the only order that keeps you honest.
So what do you do on Monday. Pick your one or two real goals and write them as sentences a user would recognise. For each, choose a signal, then a metric, in that order. Report every headline number with a percentile beside it. Keep at least one guardrail that is allowed to say no. And before any test, write the result that would change your mind, and keep the note. If you never write that note, you already know the answer you are going to find.