All concepts
Positional Encoding
Inject token order because attention alone is permutation-invariant.
Transformers & LLMs · Intermediate · ~8 min
In plain English
Attention alone can't tell word order — it sees a bag of words. Positional encoding stamps each token with where it sits so the sentence has a sequence again.
Why it's worth your time
Without it 'dog bites man' and 'man bites dog' are literally the same input, and how you encode position decides how far the model can extrapolate.
If you remember three things
- Attention is permutation-invariant on its own
- Sinusoidal is fixed; learned is trained; rotary is now standard
- How you encode position sets your extrapolation behaviour
Overview
Self-attention is permutation-invariant — it sees a set of token vectors with no inherent order. Positional encoding injects each token's position, via fixed sinusoids or learned vectors added to the embeddings, so the model can distinguish 'dog bites man' from 'man bites dog'.
How it works
- Start: Token Embeddings Self-attention sees a bag of token vectors unless position is added.
- Token Embeddings -> Position Signal Use learned vectors or sinusoidal functions to encode token index.
- Position Signal -> Add to Tokens Token identity and position are combined before attention.
- Add to Tokens -> Order-Aware The model can distinguish dog bites man from man bites dog.
In an interview
Because attention treats its inputs as a bag of vectors, you must add position information. Classic approaches add a sinusoidal or learned positional vector to each token embedding before the first layer. Sinusoids generalize to unseen lengths, while learned embeddings are simpler but capped at the maximum trained length.
Production defaults
- Modern default
- RoPE. Encodes relative position and extends beyond training length better than learned absolute
- Learned absolute
- hard-caps you at the trained length — extending means interpolating the table
- Long context
- position interpolation or NTK-aware scaling, then a short fine-tune to adapt
What breaks
- Quality falls off a cliff past the trained length — Positions never seen in training. Interpolate rather than extrapolate, and fine-tune briefly at the new length.
- Word order seems ignored — Positional information isn't reaching the model — check it's actually added/applied, not silently dropped.