All concepts

Hypothesis Testing

Test a claim by asking how surprising your data is if nothing were going on.

Maths · Advanced · ~4 min

In plain English

Assume nothing interesting happened. Then ask: if that were true, how surprising is what I actually saw? If it's surprising enough, stop assuming.

Why it's worth your time

Every A/B test, every 'is the new model better' decision, rests on this — and on not making the four classic mistakes.

If you remember three things

  • p-value = probability of data this extreme IF the null is true
  • It is NOT the probability the null is true
  • Statistical significance is not practical significance

Overview

Hypothesis testing frames a question as a null hypothesis (no effect) versus an alternative, then measures how surprising the observed data would be if the null were true. That surprise is the p-value; if it falls below a pre-set significance level α, you reject the null. The Central Limit Theorem is what makes the sampling math work.

In an interview

You assume nothing's happening (the null), collect a sample, and compute a p-value: the probability of data this extreme if the null were true. If p is below your threshold α (often 0.05), you reject the null. Crucially, p is not the probability the null is true, and a non-significant result isn't proof there's no effect.

Production defaults

Decide first
the metric, the sample size and the stopping rule, before you start. Written down
Power
80% at your minimum effect of interest. Underpowered tests mostly produce noise you'll act on
Multiple tests
correct for them (Bonferroni, Benjamini–Hochberg). Twenty metrics guarantee one 'significant' result at p<0.05

What breaks

  • Result significant on Tuesday, gone by Friday — Peeking. Checking repeatedly inflates the false-positive rate badly — use a fixed horizon or a sequential test designed for it.
  • p = 0.03 but the lift is 0.02% — Significant and irrelevant. Report the effect size and its confidence interval, not just the p-value.

Watch it explained

Idea behind hypothesis testing — Khan Academy, 9:58

Related