All concepts

Layer Norm

Normalize each token representation to stabilize deep network training.

Transformers & LLMs · Intermediate · ~8 min

In plain English

After every layer, re-centre and re-scale the numbers so they stay in a sane range — like normalizing volume between tracks so nothing blows out the speakers.

Why it's worth your time

It's the reason deep transformers train at all, and where it sits (before or after the block) changes whether training is stable.

If you remember three things

  • Normalizes across features within one example, not across the batch
  • Batch-size independent — that's why sequences use it
  • Pre-norm trains stably; post-norm needs careful warmup

Overview

LayerNorm normalizes each token's hidden vector across its feature dimension to zero mean and unit variance, then applies a learned scale γ and shift β: LN(x)=γ·(x−μ)/√(σ²+ε)+β. Unlike BatchNorm it has no batch or sequence dependence, so it stabilizes activations and gradients in deep transformers.

How it works

  1. Start: Token Vector A token's hidden vector can have drifting scale across layers.
  2. Token Vector -> Mean + Variance LayerNorm computes statistics across the feature dimension for that token.
  3. Mean + Variance -> Normalize Subtract mean and divide by standard deviation, then apply learned scale and bias.
  4. Normalize -> Stable Block Pre-norm transformers keep gradients and activations controlled through many layers.

In an interview

LayerNorm standardizes a single example's activations over the feature axis — subtract the mean, divide by √(var+ε), then rescale with learned γ and β. It removes internal scale drift with no batch dependence, which is why transformers use it. Many modern LLMs swap in the cheaper RMSNorm variant.

Production defaults

Placement
pre-norm (normalize before the sublayer). Every modern LLM does this
Variant
RMSNorm — drops the mean subtraction, slightly faster, same quality. Now the common choice
Epsilon
1e-5 to 1e-6. Too small and you divide by near-zero in fp16

What breaks

  • Deep model diverges early in training — Post-norm without enough warmup. Switch to pre-norm.
  • NaNs in fp16 — Epsilon too small or normalization computed in half precision. Compute norms in fp32.

Watch it explained

What is Layer Normalization? | Deep Learning Fundamentals — AssemblyAI, 5:18

Related