All concepts
Optimizers
Convert gradients into parameter updates using momentum, adaptive scaling, and schedules.
ML Foundations · Intermediate · ~8 min
In plain English
The gradient tells you which way is downhill. The optimizer decides how big a step to take, whether to keep momentum from the last step, and whether some directions deserve smaller steps than others.
Why it's worth your time
Swapping the optimizer or its schedule often buys more than architecture changes, and costs one line.
If you remember three things
- Momentum carries you through narrow ravines
- Adam gives every parameter its own effective step size
- The schedule matters as much as the optimizer
Overview
Algorithms that turn per-parameter gradients into weight updates, θ ← θ − α·update(g). Beyond plain SGD they add momentum to smooth noisy directions and adaptive per-parameter step sizes, plus learning-rate schedules for stable convergence.
How it works
- Start: Gradient Backprop produces a slope for every parameter.
- Gradient -> Momentum Momentum smooths noisy gradients so updates keep useful direction through ravines.
- Momentum -> Adam Adam tracks first and second moments, adapting the step size per parameter.
- Adam -> LR Schedule Warmup, cosine decay, and clipping control stability at scale.
- LR Schedule -> Updated Weights The optimizer step changes weights while trying to reduce validation loss, not just train loss.
In an interview
An optimizer decides how to step on the loss surface given gradients. SGD with momentum accumulates a velocity to push through ravines; Adam additionally tracks first and second gradient moments to adapt the step size per parameter. Schedules like warmup and cosine decay, plus gradient clipping, keep training stable at scale.
Production defaults
- Default
- AdamW, β₁=0.9, β₂=0.999 (0.95 for large-scale LLM training), weight decay 0.01–0.1
- LR
- 1e-3 training from scratch, 1e-4 to 2e-5 fine-tuning. Warmup 3–5% of steps then cosine decay
- When to use SGD
- vision models with a long training budget — well-tuned SGD+momentum still generalizes slightly better
What breaks
- Training diverges in the first 100 steps — No warmup. Adam's second-moment estimate is unreliable early; warmup exists exactly for this.
- Adam trains fast but generalizes worse than expected — Known trade-off. Use decoupled weight decay (AdamW, not Adam+L2) and lower the peak LR.