All concepts

Optimizers

Convert gradients into parameter updates using momentum, adaptive scaling, and schedules.

ML Foundations · Intermediate · ~8 min

In plain English

The gradient tells you which way is downhill. The optimizer decides how big a step to take, whether to keep momentum from the last step, and whether some directions deserve smaller steps than others.

Why it's worth your time

Swapping the optimizer or its schedule often buys more than architecture changes, and costs one line.

If you remember three things

  • Momentum carries you through narrow ravines
  • Adam gives every parameter its own effective step size
  • The schedule matters as much as the optimizer

Overview

Algorithms that turn per-parameter gradients into weight updates, θ ← θ − α·update(g). Beyond plain SGD they add momentum to smooth noisy directions and adaptive per-parameter step sizes, plus learning-rate schedules for stable convergence.

How it works

  1. Start: Gradient Backprop produces a slope for every parameter.
  2. Gradient -> Momentum Momentum smooths noisy gradients so updates keep useful direction through ravines.
  3. Momentum -> Adam Adam tracks first and second moments, adapting the step size per parameter.
  4. Adam -> LR Schedule Warmup, cosine decay, and clipping control stability at scale.
  5. LR Schedule -> Updated Weights The optimizer step changes weights while trying to reduce validation loss, not just train loss.

In an interview

An optimizer decides how to step on the loss surface given gradients. SGD with momentum accumulates a velocity to push through ravines; Adam additionally tracks first and second gradient moments to adapt the step size per parameter. Schedules like warmup and cosine decay, plus gradient clipping, keep training stable at scale.

Production defaults

Default
AdamW, β₁=0.9, β₂=0.999 (0.95 for large-scale LLM training), weight decay 0.01–0.1
LR
1e-3 training from scratch, 1e-4 to 2e-5 fine-tuning. Warmup 3–5% of steps then cosine decay
When to use SGD
vision models with a long training budget — well-tuned SGD+momentum still generalizes slightly better

What breaks

  • Training diverges in the first 100 steps — No warmup. Adam's second-moment estimate is unreliable early; warmup exists exactly for this.
  • Adam trains fast but generalizes worse than expected — Known trade-off. Use decoupled weight decay (AdamW, not Adam+L2) and lower the peak LR.

Watch it explained

Adam Optimizer Explained in Detail | Deep Learning — Learn With Jay, 5:05

Related