All concepts

Transformer Block

Attention + feed-forward, each wrapped in residual connections and normalization.

Transformers & LLMs · Intermediate · ~8 min

In plain English

One repeating unit: let the tokens talk to each other (attention), then let each token think on its own (feed-forward), with a bypass road around both.

Why it's worth your time

Stack this block N times and you have GPT, BERT, Llama and everything else — the differences are details on top of this shape.

If you remember three things

  • Attention mixes across positions; the MLP processes each position alone
  • Residual connections let gradients skip the whole block
  • Two-thirds of the parameters live in the feed-forward layers

Overview

A transformer block stacks multi-head self-attention and a position-wise feed-forward network, each wrapped with residual connections and layer normalization. Stacking many blocks builds an LLM. Residuals keep gradients flowing; normalization stabilizes training.

How it works

  1. Tokens enter the block Token vectors enter and go into multi-head self-attention.
  2. Attention mixes tokens Attention lets each token gather context from every other token.
  3. Residual + normalization A residual adds the input back and normalization stabilizes it — this is what keeps deep stacks trainable.
  4. Feed-forward transforms A position-wise FFN (usually 4× wider) transforms each token independently — most parameters live here.
  5. Residual + norm, then next block Another residual + norm, and the result flows into the next identical block. Stack N of them to build an LLM.

In an interview

A transformer block is multi-head attention followed by a feed-forward network, each with a residual connection and layer norm. Attention mixes information across tokens; the FFN transforms each token; residuals and norm keep deep stacks trainable. Stack N of them to build an LLM.

Production defaults

Shape
FFN hidden dimension ≈ 4× model dimension (or ~2.7× with SwiGLU's three matrices)
Order
pre-norm → attention → residual → pre-norm → MLP → residual
Dropout
0.0–0.1. Large-scale pretraining often uses none at all

What breaks

  • Deep stack won't converge — Missing residuals or post-norm placement. Both show up as the depth increasing and quality falling.
  • Parameter count is dominated by something unexpected — It's the FFN, almost always. That's also where mixture-of-experts does its work.

Watch it explained

Inside a Transformer Block | Explained Visually (No Math) — AIChronicles_JK, 3:43

Related