Skip to content
The Product Guys
All lessons
Metrics6 min read

Why Averages Hide the Story

A flat average often means two groups moving hard in opposite directions.


You shipped the new editor. Average session length is 12 minutes, the same as before. The team concludes the change was neutral and moves on. Three months later churn: The rate at which customers stop paying or stop using the product. is up among your largest accounts and nobody connects it to the editor, because the average never flinched.

An average is a single number standing in for a whole population, and it only tells the truth when the population is roughly one thing. Product usage is almost never one thing. It is a small group of heavy users, a large group of light users, and a long tail of people who arrived once and never came back, all stacked on top of each other.

The flat line, and what is under it

100%75%50%25%0%New accountsAverageOver a year old1357911Week of the quarterWeekly active share
Three lines from one dataset. The average has not moved all quarter, so the metric review says the change did nothing. Split by cohort and it did a great deal: new accounts are up hard, accounts older than a year are leaving, and the two cancel to a straight line.

Worked example

Hypothetical: Ostara, a design review tool

Before the change, 100 accounts. 20 power accounts average 40 minutes per session, 80 casual accounts average 5 minutes. The mean is (20 times 40 plus 80 times 5) divided by 100, which is (800 plus 400) divided by 100, so 12 minutes. After the change, the power accounts find the new editor slower and drop to 25 minutes. The casual accounts find it simpler and rise to 8.75 minutes. New mean: (20 times 25 plus 80 times 8.75) divided by 100, which is (500 plus 700) divided by 100, so 12 minutes again. Identical average, and underneath it the segment that generates most of your revenue lost nearly 40 percent of their engagement while the segment that generates least gained 75 percent. The average did not fail to notice. It is arithmetically incapable of noticing.

Two things to take from those numbers. First, the mean is a weighted sum, so a big move in a small group can be cancelled by a small move in a large one. Second, the groups here are not exotic. Power users and casual users is the most ordinary split in software, and it is enough to break the number completely.

What to look at instead

The distribution
Plot the histogram before you compute anything. You are looking for whether it has one hump or two, and how long the right tail is. Thirty seconds of looking prevents most of these mistakes.
Percentiles
Median, p25, p75, p95. For latency and anything with a long tail, the p95 is the user experience people complain about, and the mean is the one that looks fine in the deck.
Segments
Split by the dimensions your business actually cares about: plan, company size, tenure, platform. Report the metric per segment, not just overall.
Cohorts
Group by when people joined and follow each group over time. This separates a real change in behaviour from a change in who is in the denominator.

That last one deserves its own warning, because it produces the most confident wrong conclusions in product analytics. Suppose your average engagement per user rises this month. That can happen because existing users engaged more. It can also happen because you paused paid acquisition, so a flood of low intent signups stopped arriving and the denominator got healthier. Same chart, opposite meanings, and only a cohort: A group of users bucketed by when they joined, tracked over time rather than blended together. view tells you which happened. A growing product has a constantly changing population, so almost any aggregate you track is contaminated by mix.

Average thinking

  • Average session length is up 4 percent
  • Mean page load is 1.2 seconds, that is fine
  • Average revenue per account grew this quarter
  • The change was neutral overall

Distribution thinking

  • Sessions rose for accounts under 20 seats and fell above it
  • p95 load is 6 seconds, and it is concentrated in the reporting screen
  • Revenue per account grew because small accounts churned out
  • The change helped one segment and hurt the one that pays

One practical habit. Before you report any aggregate, ask what would have to be true for this number to be lying to me, and then go and check that one thing. Usually it is a mix shift. Sometimes it is a handful of outlier accounts, in which case the median will disagree with the mean and you should trust the disagreement.

The uncomfortable part is that distribution thinking produces messier readouts. You will walk into a review with four segment lines instead of one clean number, and someone will ask you to summarise. Resist a little. The clean number is what let the editor regression hide for three months.

Quick check

An aggregate engagement metric rises after you pause paid acquisition. What is the likely cause?

The takeaway

Averages collapse opposing movements into a flat line, so look at the distribution and split by cohort before you believe any aggregate.

Try this tomorrow

Take the headline metric you reported last week and re-cut it by account size and by signup month. Bring whichever cut disagrees with the average to your next review.

Answer the check above, then bank the day.

Where this comes from

  • Trustworthy Online Controlled Experiments, Ron Kohavi, Diane Tang and Ya Xu
  • The North Star Playbook, Amplitude

Metrics is one of six tracks. These lessons summarise and build on the work above, they do not reproduce it. Buy the books, they are better.