All concepts

Vanishing Gradient

Gradients shrink as they flow backward, so early layers barely learn.

ML Foundations · Intermediate · ~8 min

In plain English

A message whispered down a line of fifty people. Each person passes on slightly less than they heard, so by the end the first person's words are gone entirely.

Why it's worth your time

It's the reason deep learning didn't work for twenty years, and the reason residual connections are in literally every modern architecture.

If you remember three things

  • Backprop multiplies derivatives; many below 1 collapse the product
  • Sigmoid saturates at slope ≤ 0.25 — that's the culprit
  • Residual connections give the gradient a road with no multiplication

Overview

In deep or recurrent nets, backprop multiplies many local derivatives via the chain rule; when most are below 1, their product shrinks toward zero, so early layers receive almost no learning signal. Residual connections, normalization, and non-saturating activations mitigate it.

How it works

  1. Start: Loss The error signal starts at the output layer.
  2. Loss -> Chain Rule Backprop multiplies many local derivatives together across depth or time.
  3. Chain Rule -> Tiny Product If many derivatives are below 1, the product collapses toward zero.
  4. Tiny Product -> Early Layers Near-zero gradients mean early features update slowly or not at all.
  5. Early Layers -> Fixes Residual connections, normalization, ReLU/GELU, gating, and careful initialization preserve signal.

In an interview

The vanishing gradient problem is when the error signal decays as it flows backward through many layers. Because backprop chains multiplications of local derivatives, and saturating activations like sigmoid have derivatives ≤ 0.25, the product collapses and early layers barely update. Fixes include ReLU/GELU, batch/layer norm, residual skip connections, and gating.

Production defaults

Activation
ReLU or GELU. Never sigmoid/tanh in a deep hidden stack
Architecture
residual connections plus layer norm. This combination is why 100-layer networks train at all
Init
He for ReLU, Xavier for tanh — set so variance neither shrinks nor grows with depth

What breaks

  • Early layers barely change during training — Log per-layer gradient norms. If they fall off a cliff with depth, you need residuals, not a lower learning rate.
  • Loss explodes instead — The mirror problem — derivatives above 1. Clip the gradient norm and check your initialization.

Watch it explained

Exploding Gradient and Vanishing Gradient problem in deep neural network|Deep learning tutorial — Unfold Data Science, 8:41

Related