Why Averages Hide the Story
A flat average often means two groups moving hard in opposite directions.
You shipped the new editor. Average session length is 12 minutes, the same as before. The team concludes the change was neutral and moves on. Three months later churn: The rate at which customers stop paying or stop using the product. is up among your largest accounts and nobody connects it to the editor, because the average never flinched.
An average is a single number standing in for a whole population, and it only tells the truth when the population is roughly one thing. Product usage is almost never one thing. It is a small group of heavy users, a large group of light users, and a long tail of people who arrived once and never came back, all stacked on top of each other.
The flat line, and what is under it
Worked example
Hypothetical: Ostara, a design review tool
Before the change, 100 accounts. 20 power accounts average 40 minutes per session, 80 casual accounts average 5 minutes. The mean is (20 times 40 plus 80 times 5) divided by 100, which is (800 plus 400) divided by 100, so 12 minutes. After the change, the power accounts find the new editor slower and drop to 25 minutes. The casual accounts find it simpler and rise to 8.75 minutes. New mean: (20 times 25 plus 80 times 8.75) divided by 100, which is (500 plus 700) divided by 100, so 12 minutes again. Identical average, and underneath it the segment that generates most of your revenue lost nearly 40 percent of their engagement while the segment that generates least gained 75 percent. The average did not fail to notice. It is arithmetically incapable of noticing.
Two things to take from those numbers. First, the mean is a weighted sum, so a big move in a small group can be cancelled by a small move in a large one. Second, the groups here are not exotic. Power users and casual users is the most ordinary split in software, and it is enough to break the number completely.
What to look at instead
- The distribution
- Plot the histogram before you compute anything. You are looking for whether it has one hump or two, and how long the right tail is. Thirty seconds of looking prevents most of these mistakes.
- Percentiles
- Median, p25, p75, p95. For latency and anything with a long tail, the p95 is the user experience people complain about, and the mean is the one that looks fine in the deck.
- Segments
- Split by the dimensions your business actually cares about: plan, company size, tenure, platform. Report the metric per segment, not just overall.
- Cohorts
- Group by when people joined and follow each group over time. This separates a real change in behaviour from a change in who is in the denominator.
That last one deserves its own warning, because it produces the most confident wrong conclusions in product analytics. Suppose your average engagement per user rises this month. That can happen because existing users engaged more. It can also happen because you paused paid acquisition, so a flood of low intent signups stopped arriving and the denominator got healthier. Same chart, opposite meanings, and only a cohort: A group of users bucketed by when they joined, tracked over time rather than blended together. view tells you which happened. A growing product has a constantly changing population, so almost any aggregate you track is contaminated by mix.
Average thinking
- Average session length is up 4 percent
- Mean page load is 1.2 seconds, that is fine
- Average revenue per account grew this quarter
- The change was neutral overall
Distribution thinking
- Sessions rose for accounts under 20 seats and fell above it
- p95 load is 6 seconds, and it is concentrated in the reporting screen
- Revenue per account grew because small accounts churned out
- The change helped one segment and hurt the one that pays
One practical habit. Before you report any aggregate, ask what would have to be true for this number to be lying to me, and then go and check that one thing. Usually it is a mix shift. Sometimes it is a handful of outlier accounts, in which case the median will disagree with the mean and you should trust the disagreement.
The uncomfortable part is that distribution thinking produces messier readouts. You will walk into a review with four segment lines instead of one clean number, and someone will ask you to summarise. Resist a little. The clean number is what let the editor regression hide for three months.
Quick check
An aggregate engagement metric rises after you pause paid acquisition. What is the likely cause?
The takeaway
Averages collapse opposing movements into a flat line, so look at the distribution and split by cohort before you believe any aggregate.
Try this tomorrow
Take the headline metric you reported last week and re-cut it by account size and by signup month. Bring whichever cut disagrees with the average to your next review.
Answer the check above, then bank the day.
Where this comes from
- Trustworthy Online Controlled Experiments, Ron Kohavi, Diane Tang and Ya Xu
- The North Star Playbook, Amplitude
Metrics is one of six tracks. These lessons summarise and build on the work above, they do not reproduce it. Buy the books, they are better.
