Designing an A/B Test That Can Conclude
Decide the metric, the effect size and the duration before the test runs, or it will tell you nothing.
The test has been running nine days. Variant B is up 3 percent on conversion, the tool says not significant, and the launch date is Friday. Someone suggests letting it run over the weekend. Someone else notices that B is significantly up for mobile users specifically. By Monday the team ships B, and six weeks later conversion looks exactly like it did before.
Almost every bad experiment fails for the same reason: decisions that should have been made before the test started got made after the data arrived. Once you can see the numbers, every choice you make is contaminated by them, including choices that feel purely technical.
Fixed before the first data point
Fix these five before you start
- One primary metric
- Chosen in advance and written down. Everything else is a guardrail or an exploratory observation, and none of them can rescue a failed primary.
- Minimum detectable effect
- The smallest improvement that would change what you do. If a 1 percent lift would not change your decision, do not design a test that can only detect 1 percent.
- Sample size and duration
- Computed from your baseline rate and that effect size. Round up to whole weeks so you cover the full weekly cycle, and do not stop early because the line looks good.
- Segments to examine
- Named in advance. Anything you slice after seeing the results is a hypothesis for a future test, not a finding.
- The decision rule
- Write down now what you will do for each outcome, including the flat one. Most arguments after a test are really arguments about a rule nobody agreed on.
Worked example
Hypothetical: Wren, a subscription box checkout
Baseline checkout conversion is 4 percent, and the team gets about 5,000 checkout starts a week per variant. They want to detect a relative lift of 5 percent, meaning 4 percent to 4.2 percent. That is a small absolute difference of 0.2 percentage points on a noisy binary outcome, and the sample needed to detect it reliably runs to the tens of thousands per variant, so the test needs several weeks rather than the one week that was on the plan. That calculation is the useful output, and it should happen before anyone writes code. It gives the team three honest options: run it for the full duration, aim for a bigger change that a smaller sample can detect, or accept that this particular question is not answerable at their traffic and decide it on judgment instead. What they cannot do is run it for one week and read the result.
Underpowered tests are worse than no tests, and the reason is not obvious. If a test is underpowered, the only results that clear significance are large ones, and with a small true effect the large observed results are mostly noise. So the winners you do find are inflated, which is why so many shipped wins fail to show up in the quarterly number.
Cannot conclude
- Peeking daily and stopping on a good day
- Primary metric chosen after the results
- Segments sliced until one is significant
- Test run for four days
- Both variants also changed pricing mid test
Can conclude
- Fixed duration agreed in advance
- One primary metric, guardrails declared
- Segments pre registered
- Whole weeks, covering the weekly cycle
- One change at a time, or a factorial design
Run an A/A test before you trust the platform. Split traffic between two identical experiences and confirm you get no difference. Teams that do this often discover a bucketing bug, a logging gap, or a systematic difference between the groups. Finding that during an A/A test is free. Finding it after you have shipped three winners is not.
And hold onto the null result. A test that shows no difference is a real finding: it means the thing you were arguing about for two weeks does not matter as much as you thought, and you can stop spending on it. Kohavi and colleagues have written about how large a share of well designed experiments at mature companies fail to move the target metric. Expect that. A team whose tests always win is not testing, it is confirming.
Quick check
Why are the significant wins from an underpowered test often inflated?
The takeaway
An experiment concludes only if the metric, effect size, duration and decision rule were fixed before the first data point arrived.
Try this tomorrow
Before your next test, write a one page pre registration: primary metric, minimum detectable effect, required sample, end date, and what you will do for each outcome.
Answer the check above, then bank the day.
Where this comes from
- Trustworthy Online Controlled Experiments, Ron Kohavi, Diane Tang and Ya Xu
- Hacking Growth, Sean Ellis and Morgan Brown
Metrics is one of six tracks. These lessons summarise and build on the work above, they do not reproduce it. Buy the books, they are better.
