All concepts
Vanishing Gradient
Gradients shrink as they flow backward, so early layers barely learn.
ML Foundations · Intermediate · ~8 min
In plain English
A message whispered down a line of fifty people. Each person passes on slightly less than they heard, so by the end the first person's words are gone entirely.
Why it's worth your time
It's the reason deep learning didn't work for twenty years, and the reason residual connections are in literally every modern architecture.
If you remember three things
- Backprop multiplies derivatives; many below 1 collapse the product
- Sigmoid saturates at slope ≤ 0.25 — that's the culprit
- Residual connections give the gradient a road with no multiplication
Overview
In deep or recurrent nets, backprop multiplies many local derivatives via the chain rule; when most are below 1, their product shrinks toward zero, so early layers receive almost no learning signal. Residual connections, normalization, and non-saturating activations mitigate it.
How it works
- Start: Loss The error signal starts at the output layer.
- Loss -> Chain Rule Backprop multiplies many local derivatives together across depth or time.
- Chain Rule -> Tiny Product If many derivatives are below 1, the product collapses toward zero.
- Tiny Product -> Early Layers Near-zero gradients mean early features update slowly or not at all.
- Early Layers -> Fixes Residual connections, normalization, ReLU/GELU, gating, and careful initialization preserve signal.
In an interview
The vanishing gradient problem is when the error signal decays as it flows backward through many layers. Because backprop chains multiplications of local derivatives, and saturating activations like sigmoid have derivatives ≤ 0.25, the product collapses and early layers barely update. Fixes include ReLU/GELU, batch/layer norm, residual skip connections, and gating.
Production defaults
- Activation
- ReLU or GELU. Never sigmoid/tanh in a deep hidden stack
- Architecture
- residual connections plus layer norm. This combination is why 100-layer networks train at all
- Init
- He for ReLU, Xavier for tanh — set so variance neither shrinks nor grows with depth
What breaks
- Early layers barely change during training — Log per-layer gradient norms. If they fall off a cliff with depth, you need residuals, not a lower learning rate.
- Loss explodes instead — The mirror problem — derivatives above 1. Clip the gradient norm and check your initialization.