Test a claim by asking how surprising your data is if nothing were going on.
Assume nothing interesting happened. Then ask: if that were true, how surprising is what I actually saw? If it's surprising enough, stop assuming.
Every A/B test, every 'is the new model better' decision, rests on this — and on not making the four classic mistakes.
Hypothesis testing frames a question as a null hypothesis (no effect) versus an alternative, then measures how surprising the observed data would be if the null were true. That surprise is the p-value; if it falls below a pre-set significance level α, you reject the null. The Central Limit Theorem is what makes the sampling math work.
You assume nothing's happening (the null), collect a sample, and compute a p-value: the probability of data this extreme if the null were true. If p is below your threshold α (often 0.05), you reject the null. Crucially, p is not the probability the null is true, and a non-significant result isn't proof there's no effect.
Idea behind hypothesis testing — Khan Academy, 9:58