All concepts
Transformer Block
Attention + feed-forward, each wrapped in residual connections and normalization.
Transformers & LLMs · Intermediate · ~8 min
In plain English
One repeating unit: let the tokens talk to each other (attention), then let each token think on its own (feed-forward), with a bypass road around both.
Why it's worth your time
Stack this block N times and you have GPT, BERT, Llama and everything else — the differences are details on top of this shape.
If you remember three things
- Attention mixes across positions; the MLP processes each position alone
- Residual connections let gradients skip the whole block
- Two-thirds of the parameters live in the feed-forward layers
Overview
A transformer block stacks multi-head self-attention and a position-wise feed-forward network, each wrapped with residual connections and layer normalization. Stacking many blocks builds an LLM. Residuals keep gradients flowing; normalization stabilizes training.
How it works
- Tokens enter the block Token vectors enter and go into multi-head self-attention.
- Attention mixes tokens Attention lets each token gather context from every other token.
- Residual + normalization A residual adds the input back and normalization stabilizes it — this is what keeps deep stacks trainable.
- Feed-forward transforms A position-wise FFN (usually 4× wider) transforms each token independently — most parameters live here.
- Residual + norm, then next block Another residual + norm, and the result flows into the next identical block. Stack N of them to build an LLM.
In an interview
A transformer block is multi-head attention followed by a feed-forward network, each with a residual connection and layer norm. Attention mixes information across tokens; the FFN transforms each token; residuals and norm keep deep stacks trainable. Stack N of them to build an LLM.
Production defaults
- Shape
- FFN hidden dimension ≈ 4× model dimension (or ~2.7× with SwiGLU's three matrices)
- Order
- pre-norm → attention → residual → pre-norm → MLP → residual
- Dropout
- 0.0–0.1. Large-scale pretraining often uses none at all
What breaks
- Deep stack won't converge — Missing residuals or post-norm placement. Both show up as the depth increasing and quality falling.
- Parameter count is dominated by something unexpected — It's the FFN, almost always. That's also where mixture-of-experts does its work.