All concepts
Bayes' Theorem
Update a prior belief with evidence to get a posterior probability.
Maths · Intermediate · ~4 min
In plain English
You had a belief. New evidence arrives. Bayes tells you exactly how much to move — and the answer depends heavily on how common the thing was to begin with.
Why it's worth your time
It's the reason a 99%-accurate test for a rare disease still mostly returns false positives, and it comes up in nearly every quantitative interview.
If you remember three things
- Posterior ∝ likelihood × prior
- The base rate is what everyone forgets
- Rare event + good test can still mean most positives are false
Overview
Bayes' theorem updates a prior belief into a posterior once evidence arrives. It multiplies the prior by the likelihood of the evidence and normalizes by the evidence's total probability. The classic medical-test example shows why a rare base rate keeps the posterior modest even after a positive result.
How it works
- A population to screen Picture 1,000 people as a 100-square grid — each square is 10 people. We'll run a medical test across everyone and track who ends up positive.
- The prior: who is truly sick Only the base-rate fraction is actually sick before any test runs. This prior — how rare the condition is — is the number everyone forgets.
- Sensitivity catches the sick A good test flags most truly sick people with a '+'. Sensitivity is P(+ | sick) — the true-positive rate.
- False alarms from the healthy The healthy pool is huge, so even a small false-positive rate produces many '+' marks from people who are perfectly fine.
- You tested positive Condition on the evidence: throw away every non-'+' square. Only the positives remain — true positives and false positives mixed together.
- Posterior = sick + ÷ all + Of all the '+' squares, what fraction is genuinely sick? With a rare disease that ratio is far below the test's accuracy — often near a coin flip.
- Prior × likelihood ÷ evidence That's Bayes' theorem: the prior did most of the work. Change the base rate and the posterior swings — try the base-rate control.
In an interview
Bayes' theorem turns P(B|A) into P(A|B) by weighting with the prior and dividing by the overall evidence: P(A|B)=P(B|A)P(A)/P(B). Its big lesson is the base rate — when a condition is rare, even an accurate test yields many false positives, so a positive result may still mean only a coin-flip chance of being affected.
Production defaults
- Always ask
- what's the base rate? Without it a test's accuracy tells you almost nothing
- Work in counts
- imagine 10,000 people and fill in the four cells. Far less error-prone than the formula
- In production
- the base rate drifts. A threshold tuned last year is tuned for last year's prevalence
What breaks
- Your rare-event detector is drowning in false positives — Base-rate arithmetic, not a model bug. At 0.1% prevalence even 99% specificity yields mostly false alarms.
- Precision fell but the model didn't change — Prevalence dropped. Precision depends on the base rate; recall doesn't.