How 4k-trained models came to read 200k tokens — interpolate the positions, and know that 'fits' is not 'attends'.
A tape measure marked to one metre. To measure three metres you don't invent new numbers past the end — you rescale what's printed on it.
It explains both how 200k-token windows exist and why 'it fits in the context' is not the same claim as 'the model will use it'.
A model trained at 4k tokens fails catastrophically at 8k, because RoPE's rotation angles at unseen positions are out of distribution. Position Interpolation fixes it by squeezing positions into the trained range instead of extending past it — divide every position index by a scale factor and the angles stay in distribution, at the cost of resolution between nearby tokens. NTK-aware scaling and YaRN refine this by scaling frequency bands differently: high-frequency dimensions, which encode local order, are left almost untouched, while low-frequency ones are interpolated. A short fine-tune then adapts the model. Separately, attention sinks explain why streaming works: models dump excess attention on the first few tokens, so evicting them destroys the distribution — keep them and slide the rest.
Long-context scaling is mostly about positional encodings. A model trained at 4k sees out-of-distribution RoPE angles beyond that, so Position Interpolation compresses positions into the trained range instead; NTK-aware scaling and YaRN improve on it by treating high- and low-frequency dimensions differently, followed by a short fine-tune. Attention sinks are the separate insight behind streaming — keep the first few tokens or the attention distribution collapses.
What Is a Context Window? Long Context in LLMs Explained — CR Labs, 6:13