All concepts

Transformer Architecture

The full encoder–decoder: embed, attend, add & norm, cross-attend, then predict the next token.

Transformers & LLMs · Intermediate · ~6 min

In plain English

The whole machine: text becomes tokens, tokens become vectors, dozens of identical blocks refine them, and a final layer turns the result back into a guess at the next word.

Why it's worth your time

It's the map that makes every other transformer concept on this site fit somewhere.

If you remember three things

  • Encoder-only understands (BERT), decoder-only generates (GPT)
  • The stack is uniform — depth is repetition, not new machinery
  • The output head is a projection back to vocabulary size

Overview

The transformer maps a source sequence to a target sequence using only attention and feed-forward layers — no recurrence, so it trains in parallel. A stack of encoder blocks (self-attention + feed-forward, each wrapped in residual Add & Norm) turns the source into context-rich vectors. A stack of decoder blocks then generates the target one token at a time: masked self-attention over what it has produced, cross-attention into the encoder's output, feed-forward, and finally a linear projection + softmax over the vocabulary.

How it works

  1. Embed the source Source tokens become vectors, and positional encodings are added so word order is preserved.
  2. Encoder self-attention Every source token attends to every other, mixing context across the whole sentence in parallel.
  3. Add & Norm A residual connection adds the input back, then layer-norm keeps the signal stable through depth.
  4. Feed-forward A position-wise feed-forward network transforms each token independently, adding capacity.
  5. Encoder output (×N) Stack N identical blocks; the encoder emits a context-rich representation of the whole source.
  6. Embed the target The target tokens generated so far are embedded and get their own positional encodings.
  7. Masked self-attention The decoder attends only to earlier target tokens — the causal mask hides the future.
  8. Cross-attention The decoder's queries attend to the encoder's keys and values, reading the source context.
  9. Feed-forward + norm Another feed-forward and Add & Norm, stacked N times, refine the decoder's state.
  10. Project & predict A linear layer plus softmax over the vocabulary turns the state into next-token probabilities.

In an interview

A transformer is an encoder–decoder built entirely on attention. The encoder turns the source into context vectors with stacked self-attention + feed-forward blocks, each wrapped in residual Add & Norm. The decoder generates the target token by token using masked self-attention, then cross-attention into the encoder output, then feed-forward, and a final linear + softmax. Attention lets every position exchange information in parallel, replacing recurrence.

Production defaults

Choose
decoder-only for generation; encoder-only for classification and embeddings; encoder-decoder for strict input→output mapping
Scaling
depth, width and data scale together. Making one much larger than the others wastes the budget
Practically
you will almost never train one. Knowing the shape is for choosing and debugging, not building

What breaks

  • Using a decoder-only model for embeddings — Works poorly out of the box — causal masking means the last token hasn't seen the future. Use an embedding model.
  • Fine-tuning changed unrelated behaviour — Every layer is shared across every task. That coupling is inherent — evaluate broadly, not just on your task.

Watch it explained

Transformer Explainer- Learn About Transformer With Visualization — Krish Naik, 6:49

Related