All concepts

p-values & Significance

A p-value is the probability of seeing a difference this big if the change did nothing — not the probability that the change did nothing.

Experimentation · Intermediate · ~6 min

In plain English

Flip a fair coin ten times and getting eight heads isn't proof it's rigged — it's how often a fair coin does that anyway. The p-value is that 'how often'.

Why it's worth your time

Reading it backwards — as the chance the feature works — is the single most common statistical error in industry.

If you remember three things

  • P(data this extreme | no effect), not P(no effect | data)
  • Twenty metrics at α=0.05 finds a false winner ~64% of the time
  • Significant ≠ important; the interval carries the size

Overview

Run a test where the variants are genuinely identical and you will still measure a difference, because samples vary. The p-value quantifies that: assuming no real effect, how often would random noise produce a gap at least as large as the one observed? A small p-value means the data are surprising under 'no effect', which is evidence — but it is evidence about the data given a hypothesis, not about the hypothesis given the data. Nearly every misuse in industry comes from swapping those two, which is also why 'p = 0.04' does not mean 'a 96% chance the feature works'.

In an interview

A p-value is P(data this extreme | no real effect). It is not the probability the null is true, and 1 − p is not the probability your variant wins. The 0.05 threshold is a convention that fixes the false-positive rate at 5% per test, which is why testing twenty metrics finds a false winner roughly two-thirds of the time. Report the confidence interval — it carries the effect size, which the p-value throws away.

Production defaults

Primary metric
declare one; correct the rest (Benjamini-Hochberg for large sets)
Continuous monitoring
use a sequential test, not a fixed-horizon p-value read daily
Threshold
write the practical significance bar before the test

What breaks

  • 'p = 0.04, so 96% chance it works' — That's the transposed conditional. Report the confidence interval instead.
  • Winners never replicate — Peeking or multiplicity. Fix the stopping rule and correct for the metric count.

Watch it explained

P-values and significance tests | AP Statistics | Khan Academy — Khan Academy, 7:58

Related