All concepts
Regularization
Add a penalty to discourage overly complex models that memorize noise.
ML Foundations · Beginner · ~8 min
In plain English
You charge the model rent for complexity. It can still buy a complicated answer, but only if the extra accuracy is worth the price.
Why it's worth your time
It's the cheapest, most reliable way to close a train/validation gap — usually one parameter away from a meaningfully better model.
If you remember three things
- L2 shrinks weights smoothly; L1 drives many to exactly zero
- λ too high underfits, too low does nothing — tune on validation
- Dropout and early stopping are regularization too
Overview
Adds a penalty on model complexity to the training loss so the optimizer prefers simpler functions over ones that memorize noise. It trades a little training accuracy for better generalization, shrinking the gap between train and validation error.
How it works
- Start: Training Loss A flexible model can drive training error down by fitting noise and outliers.
- Training Loss -> Penalty L1, L2, dropout, and early stopping add a cost for complexity.
- Penalty -> Smaller Weights The optimizer now prefers simpler parameter settings unless complexity clearly improves the data loss.
- Smaller Weights -> Validation Gain Generalization improves because the learned function is smoother and less brittle.
In an interview
Regularization augments the data loss with a complexity term, L_total = L_data + λ·Ω(θ), so fitting noise now costs something. L2 shrinks weights toward zero, L1 drives many exactly to zero for sparsity, and dropout and early stopping act similarly. λ controls the strength and is tuned on validation.
Production defaults
- Search
- log-scale λ from 1e-4 to 1e1, pick the validation minimum, then check the curve isn't flat
- Which penalty
- L2 by default; L1 when you want feature selection; elastic net when features are correlated AND you want sparsity
- Deep nets
- weight decay 0.01–0.1 with AdamW, dropout 0.1 in transformers
What breaks
- Regularization made train AND validation worse — λ is too high — you're underfitting. The U-curve on validation is the thing to look at, not train loss.
- L1 zeroed out a feature you know matters — Correlated features: L1 picks one arbitrarily. Use elastic net if you need both.