All concepts
Layer Norm
Normalize each token representation to stabilize deep network training.
Transformers & LLMs · Intermediate · ~8 min
In plain English
After every layer, re-centre and re-scale the numbers so they stay in a sane range — like normalizing volume between tracks so nothing blows out the speakers.
Why it's worth your time
It's the reason deep transformers train at all, and where it sits (before or after the block) changes whether training is stable.
If you remember three things
- Normalizes across features within one example, not across the batch
- Batch-size independent — that's why sequences use it
- Pre-norm trains stably; post-norm needs careful warmup
Overview
LayerNorm normalizes each token's hidden vector across its feature dimension to zero mean and unit variance, then applies a learned scale γ and shift β: LN(x)=γ·(x−μ)/√(σ²+ε)+β. Unlike BatchNorm it has no batch or sequence dependence, so it stabilizes activations and gradients in deep transformers.
How it works
- Start: Token Vector A token's hidden vector can have drifting scale across layers.
- Token Vector -> Mean + Variance LayerNorm computes statistics across the feature dimension for that token.
- Mean + Variance -> Normalize Subtract mean and divide by standard deviation, then apply learned scale and bias.
- Normalize -> Stable Block Pre-norm transformers keep gradients and activations controlled through many layers.
In an interview
LayerNorm standardizes a single example's activations over the feature axis — subtract the mean, divide by √(var+ε), then rescale with learned γ and β. It removes internal scale drift with no batch dependence, which is why transformers use it. Many modern LLMs swap in the cheaper RMSNorm variant.
Production defaults
- Placement
- pre-norm (normalize before the sublayer). Every modern LLM does this
- Variant
- RMSNorm — drops the mean subtraction, slightly faster, same quality. Now the common choice
- Epsilon
- 1e-5 to 1e-6. Too small and you divide by near-zero in fp16
What breaks
- Deep model diverges early in training — Post-norm without enough warmup. Switch to pre-norm.
- NaNs in fp16 — Epsilon too small or normalization computed in half precision. Compute norms in fp32.