All concepts

Power & Sample Size

Work out how many users you need before you start, or you'll spend three weeks proving nothing and call it a null result.

Experimentation · Intermediate · ~5 min

In plain English

Deciding how many coin flips you need before you start. Ten flips can't detect a slightly biased coin, and no amount of staring at the result will change that.

Why it's worth your time

An underpowered test kills good ideas by producing a null result that gets reported as 'it doesn't work'.

If you remember three things

  • Required sample scales with 1/effect², so half the effect costs 4× the users
  • The MDE should come from economics, not optimism
  • Reaching significance early is peeking, not finishing early

Overview

Power is the probability of detecting an effect that is really there. An underpowered test is worse than no test: it produces a non-significant result that gets reported as 'no impact', and a good idea is killed by insufficient sample rather than by evidence. Four quantities are locked together — baseline rate, minimum detectable effect, significance level, and power — so fixing three determines the fourth. That is the calculation you do before launching, and its output is a required sample size and therefore a duration.

In an interview

Power is P(detect the effect | it exists), conventionally set at 80%. Sample size depends on the baseline rate, the minimum effect you care about, α and power. Halving the detectable effect costs four times the sample. Compute it before launching; an underpowered test's null result means 'we couldn't see it', which is routinely mis-reported as 'it doesn't work'.

Production defaults

Before launch
compute sample size and duration; publish both with the plan
More power, no traffic
CUPED, and pick the metric closest to the change
Null results
always state the effect size you were powered to detect

What breaks

  • Test says 'no difference' every time — Chronically underpowered. Check what MDE your traffic supports before designing the next one.
  • Powered for revenue per user, never conclusive — Its variance is dominated by a few large orders. Use conversion, or winsorise.

Watch it explained

Statistical Power, Clearly Explained!!! — StatQuest with Josh Starmer, 8:19

Related