All concepts
Self-Attention
Each token builds a query, matches it against every key, and pulls a weighted blend of values.
Transformers & LLMs · Intermediate · ~12 min
In plain English
Every word in the sentence gets to ask every other word 'how relevant are you to me?' and then rewrites itself as a blend of the answers it cared about.
Why it's worth your time
It's the single mechanism that made transformers work, and the most-asked question in any LLM interview.
If you remember three things
- Query asks, Key advertises, Value is what gets passed along
- Scores are scaled by √d then softmaxed into weights
- Every position sees every other in one step — that's the win over RNNs
Overview
Self-attention lets every token look at every other token and decide what matters. From each token we derive a query, key, and value. Scores are query·key, softmaxed into weights, then used to average the values. It's the core operation of transformers and why they model long-range context so well.
How it works
- Token embeddings Each token in “The cat sat down” starts as a vector. Self-attention will mix information across them.
- Make Q, K, V Project each token into a Query (what it looks for), a Key (what it offers), and a Value (its content).
- Score with Q·K We follow one token, “sat”. Its query dot-products with every key to score relevance. Scores are raw — they can be negative and don't sum to 1.
- Softmax the scores Softmax squashes “sat”'s row into weights that sum to 1 — big gaps become a sharp distribution. Watch the numbers morph.
- Weighted sum of V “sat” asks “who is my subject?” and reads mostly from “cat” (weight 0.66). Its new representation is that weighted blend of values.
- Output + what's also true Every token now carries context. Real transformers also add: multi-head attention (parallel patterns), causal masking in decoders, and positional encoding.
In an interview
Self-attention lets each token attend to all others. From each token we compute a query, key, and value; scores are scaled query·key, softmaxed into weights, then used to average the values. Multi-head attention does this in parallel subspaces. It's the transformer's core and models long-range dependencies in O(n²).
Production defaults
- Cost
- O(n²) in sequence length, for time AND memory. This is the constraint behind every long-context trick
- Scaling
- dividing by √d keeps the softmax out of its saturated region
- Explain it as
- differentiable dictionary lookup with soft, learned keys
What breaks
- Memory explodes on long sequences — The n² attention matrix. FlashAttention avoids materializing it; windowed attention avoids computing all of it.
- Model ignores a fact in the middle of a long prompt — Attention dilution over long contexts. Rerank so the important content sits at the ends.