All concepts
Transformer Architecture
The full encoder–decoder: embed, attend, add & norm, cross-attend, then predict the next token.
Transformers & LLMs · Intermediate · ~6 min
In plain English
The whole machine: text becomes tokens, tokens become vectors, dozens of identical blocks refine them, and a final layer turns the result back into a guess at the next word.
Why it's worth your time
It's the map that makes every other transformer concept on this site fit somewhere.
If you remember three things
- Encoder-only understands (BERT), decoder-only generates (GPT)
- The stack is uniform — depth is repetition, not new machinery
- The output head is a projection back to vocabulary size
Overview
The transformer maps a source sequence to a target sequence using only attention and feed-forward layers — no recurrence, so it trains in parallel. A stack of encoder blocks (self-attention + feed-forward, each wrapped in residual Add & Norm) turns the source into context-rich vectors. A stack of decoder blocks then generates the target one token at a time: masked self-attention over what it has produced, cross-attention into the encoder's output, feed-forward, and finally a linear projection + softmax over the vocabulary.
How it works
- Embed the source Source tokens become vectors, and positional encodings are added so word order is preserved.
- Encoder self-attention Every source token attends to every other, mixing context across the whole sentence in parallel.
- Add & Norm A residual connection adds the input back, then layer-norm keeps the signal stable through depth.
- Feed-forward A position-wise feed-forward network transforms each token independently, adding capacity.
- Encoder output (×N) Stack N identical blocks; the encoder emits a context-rich representation of the whole source.
- Embed the target The target tokens generated so far are embedded and get their own positional encodings.
- Masked self-attention The decoder attends only to earlier target tokens — the causal mask hides the future.
- Cross-attention The decoder's queries attend to the encoder's keys and values, reading the source context.
- Feed-forward + norm Another feed-forward and Add & Norm, stacked N times, refine the decoder's state.
- Project & predict A linear layer plus softmax over the vocabulary turns the state into next-token probabilities.
In an interview
A transformer is an encoder–decoder built entirely on attention. The encoder turns the source into context vectors with stacked self-attention + feed-forward blocks, each wrapped in residual Add & Norm. The decoder generates the target token by token using masked self-attention, then cross-attention into the encoder output, then feed-forward, and a final linear + softmax. Attention lets every position exchange information in parallel, replacing recurrence.
Production defaults
- Choose
- decoder-only for generation; encoder-only for classification and embeddings; encoder-decoder for strict input→output mapping
- Scaling
- depth, width and data scale together. Making one much larger than the others wastes the budget
- Practically
- you will almost never train one. Knowing the shape is for choosing and debugging, not building
What breaks
- Using a decoder-only model for embeddings — Works poorly out of the box — causal masking means the last token hasn't seen the future. Use an embedding model.
- Fine-tuning changed unrelated behaviour — Every layer is shared across every task. That coupling is inherent — evaluate broadly, not just on your task.