All concepts

Self-Attention

Each token builds a query, matches it against every key, and pulls a weighted blend of values.

Transformers & LLMs · Intermediate · ~12 min

In plain English

Every word in the sentence gets to ask every other word 'how relevant are you to me?' and then rewrites itself as a blend of the answers it cared about.

Why it's worth your time

It's the single mechanism that made transformers work, and the most-asked question in any LLM interview.

If you remember three things

  • Query asks, Key advertises, Value is what gets passed along
  • Scores are scaled by √d then softmaxed into weights
  • Every position sees every other in one step — that's the win over RNNs

Overview

Self-attention lets every token look at every other token and decide what matters. From each token we derive a query, key, and value. Scores are query·key, softmaxed into weights, then used to average the values. It's the core operation of transformers and why they model long-range context so well.

How it works

  1. Token embeddings Each token in “The cat sat down” starts as a vector. Self-attention will mix information across them.
  2. Make Q, K, V Project each token into a Query (what it looks for), a Key (what it offers), and a Value (its content).
  3. Score with Q·K We follow one token, “sat”. Its query dot-products with every key to score relevance. Scores are raw — they can be negative and don't sum to 1.
  4. Softmax the scores Softmax squashes “sat”'s row into weights that sum to 1 — big gaps become a sharp distribution. Watch the numbers morph.
  5. Weighted sum of V “sat” asks “who is my subject?” and reads mostly from “cat” (weight 0.66). Its new representation is that weighted blend of values.
  6. Output + what's also true Every token now carries context. Real transformers also add: multi-head attention (parallel patterns), causal masking in decoders, and positional encoding.

In an interview

Self-attention lets each token attend to all others. From each token we compute a query, key, and value; scores are scaled query·key, softmaxed into weights, then used to average the values. Multi-head attention does this in parallel subspaces. It's the transformer's core and models long-range dependencies in O(n²).

Production defaults

Cost
O(n²) in sequence length, for time AND memory. This is the constraint behind every long-context trick
Scaling
dividing by √d keeps the softmax out of its saturated region
Explain it as
differentiable dictionary lookup with soft, learned keys

What breaks

  • Memory explodes on long sequences — The n² attention matrix. FlashAttention avoids materializing it; windowed attention avoids computing all of it.
  • Model ignores a fact in the middle of a long prompt — Attention dilution over long contexts. Rerank so the important content sits at the ends.

Watch it explained

Attention mechanism: Overview — Google Cloud Tech, 5:33

Related