All concepts

RoPE Positional Encoding

Encode relative positions by rotating query and key vectors.

Transformers & LLMs · Advanced · ~8 min

In plain English

Instead of adding a position tag, rotate each token's vector by an angle proportional to its position. Two tokens' relative angle then encodes their distance.

Why it's worth your time

It's what nearly every current open-weights LLM uses, and the mechanism behind every context-extension trick you'll read about.

If you remember three things

  • Rotation applied to queries and keys, not to values
  • Relative position falls out of the dot product for free
  • Extends past training length far more gracefully than learned tables

Overview

RoPE encodes position by rotating query and key vectors: it splits features into pairs and rotates each pair by an angle proportional to the token's index and the pair's frequency. Because ⟨R_m q, R_n k⟩ depends only on the offset m−n, attention scores become naturally relative — with no additive term.

How it works

  1. Start: Q and K RoPE applies position-dependent rotations to query and key vectors.
  2. Q and K -> Rotate Pairs Feature pairs are rotated by angles that depend on token position and frequency.
  3. Rotate Pairs -> Relative Offset Dot products naturally depend on distance between positions.
  4. Relative Offset -> Length Extrapolation RoPE is widely used in LLMs and can be scaled for longer context windows.

In an interview

Rotary Positional Embedding multiplies each query and key by a rotation matrix whose angle depends on absolute position. The dot product of a query rotated at position m and a key at position n depends only on m−n, so it injects relative position directly into attention scores. It adds no embeddings and can be stretched to longer contexts via frequency scaling.

Production defaults

Base θ
10000 conventionally. Raising it (NTK scaling) is the standard way to stretch context
Extension
linear position interpolation or YaRN, then a short fine-tune at the target length
Don't
just raise max_position and hope. Untuned extrapolation degrades sharply

What breaks

  • Extended context gives incoherent output — Naive extrapolation. Use interpolation/YaRN and fine-tune at the new length.
  • Fine-tune at long context is unstable — Mixed RoPE settings between base and fine-tune. They must match exactly.

Watch it explained

Positional embeddings in transformers EXPLAINED | Demystifying positional encodings. — AI Coffee Break with Letitia, 9:39

Related