All concepts

Attention Masking

Block illegal attention links so tokens only see what they are allowed to see.

Transformers & LLMs · Intermediate · ~8 min

In plain English

Blindfold the model so it can't peek at words it isn't supposed to have seen yet — future tokens when generating, padding when batching.

Why it's worth your time

A missing causal mask makes training loss look wonderful and generation output garbage, and it's a classic interview trap.

If you remember three things

  • Causal mask: position i may only see positions ≤ i
  • Padding mask stops attention wasting weight on filler
  • Masking is applied before the softmax, as −∞

Overview

Forbids certain token-pair links by adding negative infinity to their raw scores before softmax, driving those probabilities to zero. It enforces causality in decoders (no peeking at future tokens) and prevents padding positions from contaminating representations.

How it works

  1. Start: Attention Scores Raw attention scores exist for every token pair.
  2. Attention Scores -> Mask Matrix A causal or padding mask sets forbidden links to negative infinity.
  3. Mask Matrix -> Softmax After softmax, masked positions get probability zero.
  4. Softmax -> Legal Context Decoders cannot peek at future tokens, and padding cannot pollute representations.

In an interview

Masking controls which tokens can attend to which. Before the softmax you add a mask matrix to the QK^T scores, setting forbidden links to -infinity so they get zero weight: softmax((QK^T + mask)/√d_k). A causal mask blocks attending to future positions for autoregressive generation; a padding mask ignores filler tokens in batched sequences.

Production defaults

Decoder
causal mask, always. Encoder: bidirectional, no causal mask
Padding
mask it AND exclude it from the loss. Two separate mistakes people make
Implementation
add −inf (not 0) pre-softmax, so masked positions get exactly zero weight

What breaks

  • Training loss near zero, generation is nonsense — The causal mask is missing — the model was copying the answer from the future during training.
  • Batched results differ from unbatched — Padding isn't masked. The model is attending to filler tokens.

Watch it explained

Causal Masking Explained: How GPT Models Prevent Cheating During Training — Puru Kathuria, 4:10

Related