Skip to content
The Product Guys
All lessons
Metrics5 min read

Goodhart's Law and Metric Gaming

Nobody has to cheat. A target quietly rewrites what everyone thinks the job is.


Support gets a target: average first response time under two hours. Within a month they hit it. Within two months, customer satisfaction drops, because the fastest way to respond in two hours is a templated reply that asks a clarifying question and restarts the clock. Nobody cheated. Everybody did what they were measured on, and the measure stopped meaning what it used to mean.

When a measure becomes a target, it ceases to be a good measure.
Marilyn Strathern, restating Goodhart's law

The target is met, the thing it stood for is not

1007550250Tickets closed per weekIssues actually resolved012345Month after the target was setChange from the baseline ( (indexed, schematic))
Nobody cheats. The team is asked for tickets closed, so ambiguous tickets get closed and reopened rather than solved, and closure climbs exactly as intended. The thing closure was standing in for, whether people got their problem fixed, quietly goes the other way.

The mechanism is worth understanding, because teams usually treat this as a problem of bad actors and respond by adding oversight. It is not about bad actors. A metric is always a compressed proxy for something richer. Response time is a proxy for customers feeling looked after. When you target the proxy hard enough, people optimise the proxy, and the gap between the proxy and the real thing is exactly where the damage collects.

How gaming shows up

Optimising the proxy
Fast, empty responses. Tickets closed and reopened. Onboarding steps marked complete by a default value. The metric improves, the thing it stood for does not.
Shifting the denominator
Improve conversion by tightening who counts as a lead. Improve retention by narrowing the definition of active. Nothing about the product changed.
Borrowing from the future
Discounts and one off pushes that hit this quarter's number and cost next quarter's. Very common near a deadline, and almost never labelled as what it is.
Starving the unmeasured
The quietest one. Reliability, accessibility, documentation and internal tooling all decay because nothing on the dashboard notices them.

Worked example

Hypothetical: Tellwood, a learning platform

Tellwood targets course completions and ships a lighter final quiz, an auto advance on video, and reminder emails. Completions rise from 22 percent to 38 percent in a quarter, which looks like a large win. But the average time a learner spends in a course drops, and the share of learners who buy a second course falls. The completion number went up because completion got easier, not because more people learned. Here is the part that should worry you: if the team had only tracked completions, this would have been reported as a triumph, and the next quarter's target would have been set higher.

Invites gaming

  • One number, targeted hard, tied to compensation
  • Definition owned by the team being measured
  • Target set by extrapolating last quarter
  • Metric reviewed only when it is missed

Resists gaming

  • Metric paired with a counter metric
  • Definition owned centrally and version controlled
  • Target tied to a stated theory of why it should move
  • Metric and its definition reviewed on a schedule

counter metric: A metric watched alongside a target to catch the damage that hitting the target might cause. are the single most effective defence. Pair speed with satisfaction, volume with quality, completion with a downstream behaviour that only happens if real value was delivered. The pair makes the shortcut visible, because most shortcuts improve one at the direct expense of the other.

The second defence is to ask the gaming question out loud during planning. If you wanted to move this number by Friday without helping a single customer, how would you do it. People answer quickly and accurately, because they can already see the shortcuts. Write the answers down, then either change the metric or add the guardrail that catches them.

Third, keep metric definitions out of the hands of the team whose performance they measure. Not from distrust, but because the pressure to make a definition slightly more convenient is constant and nobody experiences a single small change as dishonest. A version controlled definition with a visible history means a change in the number and a change in the definition can never be confused.

And treat any sudden step change in a metric as a definition change until proven otherwise. Real product improvements move numbers gradually. A clean jump on a Tuesday is usually instrumentation, a filter, or a new default.

Quick check

Why is adding oversight a weak response to metric gaming?

The takeaway

Targets pull effort toward the proxy and away from the thing it stood for, so pair every metric with a counter metric and audit its definition.

Try this tomorrow

In your next planning meeting, ask the team how they would move the headline metric by Friday without helping a customer. Add a guardrail for each answer.

Answer the check above, then bank the day.

Where this comes from

  • Measure What Matters, John Doerr
  • Radical Focus, Christina Wodtke
  • Escaping the Build Trap, Melissa Perri

Metrics is one of six tracks. These lessons summarise and build on the work above, they do not reproduce it. Buy the books, they are better.