All concepts

Long-Context Scaling

How 4k-trained models came to read 200k tokens — interpolate the positions, and know that 'fits' is not 'attends'.

Advanced LLM Systems · Advanced · ~6 min

In plain English

A tape measure marked to one metre. To measure three metres you don't invent new numbers past the end — you rescale what's printed on it.

Why it's worth your time

It explains both how 200k-token windows exist and why 'it fits in the context' is not the same claim as 'the model will use it'.

If you remember three things

  • Interpolate positions into the trained range, don't extrapolate past it
  • YaRN scales frequency bands differently — local order survives
  • Attention sinks: never evict the first tokens from a sliding cache

Overview

A model trained at 4k tokens fails catastrophically at 8k, because RoPE's rotation angles at unseen positions are out of distribution. Position Interpolation fixes it by squeezing positions into the trained range instead of extending past it — divide every position index by a scale factor and the angles stay in distribution, at the cost of resolution between nearby tokens. NTK-aware scaling and YaRN refine this by scaling frequency bands differently: high-frequency dimensions, which encode local order, are left almost untouched, while low-frequency ones are interpolated. A short fine-tune then adapts the model. Separately, attention sinks explain why streaming works: models dump excess attention on the first few tokens, so evicting them destroys the distribution — keep them and slide the rest.

In an interview

Long-context scaling is mostly about positional encodings. A model trained at 4k sees out-of-distribution RoPE angles beyond that, so Position Interpolation compresses positions into the trained range instead; NTK-aware scaling and YaRN improve on it by treating high- and low-frequency dimensions differently, followed by a short fine-tune. Attention sinks are the separate insight behind streaming — keep the first few tokens or the attention distribution collapses.

Production defaults

Scaling
YaRN over naive linear interpolation, plus a few hundred fine-tune steps at the target length
Streaming
keep 4 sink tokens pinned plus the sliding window
Verification
needle-in-a-haystack at many depths on YOUR content — advertised window length says nothing about attention

What breaks

  • Output degrades past the trained length — Naive extrapolation puts RoPE angles out of distribution. Configure rope_scaling and fine-tune briefly.
  • Streaming output collapses into repetition — You evicted the first tokens. The model parks surplus attention there; keep the sinks.

Watch it explained

What Is a Context Window? Long Context in LLMs Explained — CR Labs, 6:13

Related