All concepts
Bias–Variance Tradeoff
Too simple underfits (high bias); too flexible overfits (high variance). Generalization lives in between.
ML Foundations · Intermediate · ~8 min
In plain English
Bias is a dartboard where every throw lands together but off-target. Variance is throws scattered all around the bullseye. You need both tight and centred.
Why it's worth your time
It's the diagnosis step. Whether you add data, add features, or simplify the model depends entirely on which of the two is hurting you.
If you remember three things
- High bias: train and validation error both high and close
- High variance: train error low, validation error much higher
- More data fixes variance. It does not fix bias
Overview
Test error decomposes into bias (error from wrong assumptions), variance (sensitivity to the training sample), and irreducible noise. Simple models have high bias; complex models have high variance. The art is finding the model complexity that minimizes their sum.
How it works
- Underfit — high bias A model that's too simple (a flat line) misses the real pattern. Both train and test error are high.
- Good fit — balanced A model with the right flexibility captures the trend without chasing noise. Test error is lowest here.
- Overfit — high variance A very flexible model snakes through every point, memorizing noise. Train error is tiny but test error explodes.
- Error decomposition Expected test error = bias² + variance + irreducible noise. Complexity trades bias for variance.
- The sweet spot As complexity rises, bias falls and variance grows. The minimum of their sum is the model you want.
In an interview
Test error splits into bias, variance, and irreducible noise. Simple models underfit (high bias); complex models overfit (high variance). We tune complexity, regularization, and data size to minimize their sum — that's the tradeoff.
Production defaults
- Diagnose first
- plot train vs validation error against training-set size before touching the model
- High bias
- more features, less regularization, a more flexible model
- High variance
- more data, more regularization, fewer features, or bagging
What breaks
- You added data and nothing improved — You were bias-limited. The learning curves would have told you before you spent the labelling budget.
- Validation looks fine, production doesn't — That's a third thing — distribution shift, not bias or variance. Check your splits are realistic.