All concepts

Anatomy of an A/B Test

Split users at random, change one thing, and the difference you measure is caused by the change — that last clause is the whole point, and randomisation is what buys it.

Experimentation · Beginner · ~6 min

In plain English

Two identical queues, one gets the new till software. Because people were sent to queues at random, any difference in speed is the software.

Why it's worth your time

It's the only routinely available tool that licenses the word 'caused'.

If you remember three things

  • Randomisation balances the things you never measured — that's the whole trick
  • Pre-register the primary metric, MDE, sample size and guardrails
  • Run whole weeks; check the split before looking at outcomes

Overview

An A/B test is the only routinely available tool that gives a causal answer. Randomly assigning users to control and treatment makes the two groups equivalent in expectation on everything — including the things you never thought to measure — so any difference in outcome is attributable to the change. Everything else about running one is protecting that property: assigning consistently, deciding the sample size in advance, choosing one primary metric before you look, and running for whole weeks so that day-of-week composition matches.

In an interview

Randomise users into control and treatment, change exactly one thing, and compare a pre-declared primary metric. Randomisation makes the groups comparable on unobserved variables, which is what licenses a causal claim. The discipline is up front: fixed sample size, one primary metric, guardrails, and no peeking — because deciding when to stop after seeing the data invalidates the statistics.

Production defaults

Unit
a stable user identifier, verified consistent in the data
Duration
complete weeks only, so weekday mix matches
Reporting
confidence interval plus the effect size you were powered for

What breaks

  • Result flipped after running longer — You stopped at the first significant reading. That's peeking; the first result was noise.
  • Marketplace test shows no effect — Interference — treated and control users compete for the same supply. Randomise by market or time.

Watch it explained

Simple explanation of A/B Testing — codebasics, 5:49

Related