All concepts

Cross Attention

Let one sequence query information from another sequence.

Transformers & LLMs · Intermediate · ~8 min

In plain English

Same mechanism as self-attention, except the questions come from one sequence and the answers come from a different one — the way a translator consults the source text.

Why it's worth your time

It's how any model conditions on something external: an image, a source language, a retrieved document.

If you remember three things

  • Queries from the target, keys and values from the source
  • It's the bridge between two modalities or two sequences
  • No causal mask on the source side — you may see all of it

Overview

Lets one sequence query information from a different sequence: queries come from the target while keys and values come from a separate source. It is how decoders pull in encoder context, powering translation, captioning, and retrieval-augmented generation.

How it works

  1. Start: Decoder Query The target sequence creates queries asking what it needs next.
  2. Decoder Query -> Source Memory A separate source sequence provides keys and values.
  3. Source Memory -> Alignment The decoder query attends to relevant source positions.
  4. Alignment -> Context Mix Retrieved source content is mixed into the target representation.
  5. Context Mix -> Generated Token Cross-attention powers translation, captioning, and retrieval-augmented decoders.

In an interview

Cross-attention is attention where queries and keys/values come from different sequences. The decoder's current representation forms the queries; the encoder's outputs (or retrieved documents) supply keys and values, so the decoder aligns to and pulls in relevant source content. It is the bridge between two sequences or modalities in encoder-decoder models.

Production defaults

Use for
encoder-decoder translation, image conditioning in diffusion, adapters over frozen encoders
Cache
source keys and values are fixed per request — compute once and reuse across all decoding steps
vs RAG
cross-attention conditions inside the model; RAG conditions through the prompt. RAG is far easier to ship

What breaks

  • The model ignores the conditioning — Cross-attention weights collapsed. Check the conditioning isn't being masked out, and that its scale matches the hidden states.
  • Slow decoding — Re-encoding the source every step. Cache source K/V once per request.

Watch it explained

Cross Attention in Transformers Explained: The Bridge Between Encoder and Decoder — Skill Advancement, 6:37

Related