All concepts

Positional Encoding

Inject token order because attention alone is permutation-invariant.

Transformers & LLMs · Intermediate · ~8 min

In plain English

Attention alone can't tell word order — it sees a bag of words. Positional encoding stamps each token with where it sits so the sentence has a sequence again.

Why it's worth your time

Without it 'dog bites man' and 'man bites dog' are literally the same input, and how you encode position decides how far the model can extrapolate.

If you remember three things

  • Attention is permutation-invariant on its own
  • Sinusoidal is fixed; learned is trained; rotary is now standard
  • How you encode position sets your extrapolation behaviour

Overview

Self-attention is permutation-invariant — it sees a set of token vectors with no inherent order. Positional encoding injects each token's position, via fixed sinusoids or learned vectors added to the embeddings, so the model can distinguish 'dog bites man' from 'man bites dog'.

How it works

  1. Start: Token Embeddings Self-attention sees a bag of token vectors unless position is added.
  2. Token Embeddings -> Position Signal Use learned vectors or sinusoidal functions to encode token index.
  3. Position Signal -> Add to Tokens Token identity and position are combined before attention.
  4. Add to Tokens -> Order-Aware The model can distinguish dog bites man from man bites dog.

In an interview

Because attention treats its inputs as a bag of vectors, you must add position information. Classic approaches add a sinusoidal or learned positional vector to each token embedding before the first layer. Sinusoids generalize to unseen lengths, while learned embeddings are simpler but capped at the maximum trained length.

Production defaults

Modern default
RoPE. Encodes relative position and extends beyond training length better than learned absolute
Learned absolute
hard-caps you at the trained length — extending means interpolating the table
Long context
position interpolation or NTK-aware scaling, then a short fine-tune to adapt

What breaks

  • Quality falls off a cliff past the trained length — Positions never seen in training. Interpolate rather than extrapolate, and fine-tune briefly at the new length.
  • Word order seems ignored — Positional information isn't reaching the model — check it's actually added/applied, not silently dropped.

Watch it explained

How positional encoding works in transformers? — BrainDrain, 5:35

Related