All concepts

Regularization

Add a penalty to discourage overly complex models that memorize noise.

ML Foundations · Beginner · ~8 min

In plain English

You charge the model rent for complexity. It can still buy a complicated answer, but only if the extra accuracy is worth the price.

Why it's worth your time

It's the cheapest, most reliable way to close a train/validation gap — usually one parameter away from a meaningfully better model.

If you remember three things

  • L2 shrinks weights smoothly; L1 drives many to exactly zero
  • λ too high underfits, too low does nothing — tune on validation
  • Dropout and early stopping are regularization too

Overview

Adds a penalty on model complexity to the training loss so the optimizer prefers simpler functions over ones that memorize noise. It trades a little training accuracy for better generalization, shrinking the gap between train and validation error.

How it works

  1. Start: Training Loss A flexible model can drive training error down by fitting noise and outliers.
  2. Training Loss -> Penalty L1, L2, dropout, and early stopping add a cost for complexity.
  3. Penalty -> Smaller Weights The optimizer now prefers simpler parameter settings unless complexity clearly improves the data loss.
  4. Smaller Weights -> Validation Gain Generalization improves because the learned function is smoother and less brittle.

In an interview

Regularization augments the data loss with a complexity term, L_total = L_data + λ·Ω(θ), so fitting noise now costs something. L2 shrinks weights toward zero, L1 drives many exactly to zero for sparsity, and dropout and early stopping act similarly. λ controls the strength and is tuned on validation.

Production defaults

Search
log-scale λ from 1e-4 to 1e1, pick the validation minimum, then check the curve isn't flat
Which penalty
L2 by default; L1 when you want feature selection; elastic net when features are correlated AND you want sparsity
Deep nets
weight decay 0.01–0.1 with AdamW, dropout 0.1 in transformers

What breaks

  • Regularization made train AND validation worse — λ is too high — you're underfitting. The U-curve on validation is the thing to look at, not train loss.
  • L1 zeroed out a feature you know matters — Correlated features: L1 picks one arbitrarily. Use elastic net if you need both.

Watch it explained

2.3 | Deep Learning | Regularization L1 and L2 | KCS-078 | AKTU & Other Universities — Engineer Being, 7:28

Related