Encode relative positions by rotating query and key vectors.
Instead of adding a position tag, rotate each token's vector by an angle proportional to its position. Two tokens' relative angle then encodes their distance.
It's what nearly every current open-weights LLM uses, and the mechanism behind every context-extension trick you'll read about.
RoPE encodes position by rotating query and key vectors: it splits features into pairs and rotates each pair by an angle proportional to the token's index and the pair's frequency. Because ⟨R_m q, R_n k⟩ depends only on the offset m−n, attention scores become naturally relative — with no additive term.
Rotary Positional Embedding multiplies each query and key by a rotation matrix whose angle depends on absolute position. The dot product of a query rotated at position m and a key at position n depends only on m−n, so it injects relative position directly into attention scores. It adds no embeddings and can be stretched to longer contexts via frequency scaling.
Positional embeddings in transformers EXPLAINED | Demystifying positional encodings. — AI Coffee Break with Letitia, 9:39