All concepts
Backpropagation
Apply the chain rule backward through the network to get every weight's gradient in one efficient pass.
Deep Learning · Intermediate · ~10 min
In plain English
After a wrong answer you walk back through the workings, asking at every step 'how much did YOU contribute to that mistake?' — and nudge each part by its share of the blame.
Why it's worth your time
It's the algorithm that makes deep learning possible at all, and the one interviewers use to check whether you understand the chain rule.
If you remember three things
- Forward pass stores activations; backward pass reuses them
- It's the chain rule applied efficiently, nothing more
- Memory cost is proportional to the stored activations, not the parameters
Overview
Backpropagation computes the gradient of the loss with respect to every weight by applying the chain rule from the output backward. It reuses cached forward activations so the whole gradient costs about as much as one forward pass — the efficiency that makes deep learning trainable.
How it works
- Forward pass Run inputs through the network, caching each layer's activations. These are needed for the backward pass.
- Compute the loss Compare the output to the target with a loss function. This scalar is what we differentiate.
- Output gradient Differentiate the loss w.r.t. the output layer — the starting error signal δ that flows backward.
- Chain rule backward Propagate δ to earlier layers by multiplying by weights and the activation derivative. Each layer's gradient reuses the next layer's.
- Update weights Each weight's gradient is δ times the incoming activation. Gradient descent then steps the weights downhill.
In an interview
Backpropagation computes the gradient of the loss with respect to every parameter by applying the chain rule backward from the output, reusing cached forward activations. It's reverse-mode autodiff — the whole gradient costs about one forward pass, which is what makes training deep nets feasible.
Production defaults
- Memory
- activation checkpointing trades ~30% extra compute for a large memory saving on deep models
- Precision
- bf16 mixed precision is the standard; keep a fp32 master copy of the weights
- Debug
- gradient-check a tiny model numerically once. It catches sign and indexing errors nothing else will
What breaks
- Out of memory on the backward pass — It's the stored activations, not the weights. Reduce batch size or turn on checkpointing.
- Gradients are all zero — A detached tensor or a dead ReLU. Print gradient norms per layer — the zeros will be contiguous.