All concepts
Gradient Descent
Walk downhill on the loss surface by repeatedly stepping opposite the gradient.
ML Foundations · Beginner · ~9 min
In plain English
You're on a foggy hillside and want the valley floor. You can only feel the slope under your feet, so you take a step downhill, feel again, and repeat.
Why it's worth your time
Almost every model you'll ever train — from logistic regression to a 400-billion-parameter LLM — is trained by this one loop.
If you remember three things
- The gradient points uphill; you step the other way
- Learning rate is the whole game: too small crawls, too big diverges
- Batch size trades gradient noise against hardware efficiency
Overview
Gradient descent is the workhorse that trains almost everything, from linear regression to giant neural networks. You compute the slope of the loss with respect to each parameter, then nudge parameters in the downhill direction. The learning rate controls step size — the single most important knob.
How it works
- The loss landscape Plot loss as a function of a parameter. Our goal is the lowest point (the minimum).
- Start somewhere Initialize the parameter θ, often randomly (marked θ₀ on the axis). We're on the hillside, not yet at the bottom.
- Compute the gradient The gradient is the multi-dimensional slope of the loss landscape at our current position. It points in the direction of steepest UPHILL climb. By calculating it, we know exactly which way to go to make the error worse — and therefore, which way to go to make it better.
- Step downhill We take a step downhill by updating the parameter θ in the opposite direction of the gradient. We scale this step by the learning rate α. If α is too large, we might overshoot the valley entirely; if α is too small, the model will take forever to train.
- Converge Repeat until the gradient is ~0. On a convex loss that's the global minimum; on neural nets, a good local one.
In an interview
Gradient descent minimizes a loss by iteratively stepping parameters in the negative gradient direction, scaled by a learning rate. Variants like SGD, momentum, and Adam trade off speed, noise, and stability. The learning rate is the key hyperparameter.
Production defaults
- Optimizer
- AdamW for deep nets, LR 1e-3 (small models) to 1e-4 (fine-tuning). SGD+momentum 0.9 when you have time to tune
- Schedule
- linear warmup for the first ~5% of steps, then cosine decay. Warmup exists to stop early instability
- Clipping
- clip gradient norm at 1.0 — cheap insurance against a single bad batch wrecking a run
What breaks
- Loss becomes NaN — Learning rate too high or an exploding gradient. Halve the LR and add gradient clipping before changing anything else.
- Loss plateaus early — LR too low, or you're in a flat region. Try an LR range test rather than guessing.