All concepts

Bias–Variance Tradeoff

Too simple underfits (high bias); too flexible overfits (high variance). Generalization lives in between.

ML Foundations · Intermediate · ~8 min

In plain English

Bias is a dartboard where every throw lands together but off-target. Variance is throws scattered all around the bullseye. You need both tight and centred.

Why it's worth your time

It's the diagnosis step. Whether you add data, add features, or simplify the model depends entirely on which of the two is hurting you.

If you remember three things

  • High bias: train and validation error both high and close
  • High variance: train error low, validation error much higher
  • More data fixes variance. It does not fix bias

Overview

Test error decomposes into bias (error from wrong assumptions), variance (sensitivity to the training sample), and irreducible noise. Simple models have high bias; complex models have high variance. The art is finding the model complexity that minimizes their sum.

How it works

  1. Underfit — high bias A model that's too simple (a flat line) misses the real pattern. Both train and test error are high.
  2. Good fit — balanced A model with the right flexibility captures the trend without chasing noise. Test error is lowest here.
  3. Overfit — high variance A very flexible model snakes through every point, memorizing noise. Train error is tiny but test error explodes.
  4. Error decomposition Expected test error = bias² + variance + irreducible noise. Complexity trades bias for variance.
  5. The sweet spot As complexity rises, bias falls and variance grows. The minimum of their sum is the model you want.

In an interview

Test error splits into bias, variance, and irreducible noise. Simple models underfit (high bias); complex models overfit (high variance). We tune complexity, regularization, and data size to minimize their sum — that's the tradeoff.

Production defaults

Diagnose first
plot train vs validation error against training-set size before touching the model
High bias
more features, less regularization, a more flexible model
High variance
more data, more regularization, fewer features, or bagging

What breaks

  • You added data and nothing improved — You were bias-limited. The learning curves would have told you before you spent the labelling budget.
  • Validation looks fine, production doesn't — That's a third thing — distribution shift, not bias or variance. Check your splits are realistic.

Watch it explained

Machine Learning Fundamentals: Bias and Variance — StatQuest with Josh Starmer, 6:36

Related