All concepts
Context Window
The maximum tokens the model can read and generate in one request.
Transformers & LLMs · Beginner · ~8 min
In plain English
The model's desk. Everything it can see this turn has to fit on it — instructions, history, retrieved documents, and room to write the answer.
Why it's worth your time
Huge windows made people stop budgeting, and both quality and cost quietly got worse. The window is a budget.
If you remember three things
- Attention cost grows with the square of the length
- Facts in the middle of a long context are the ones missed
- Output tokens compete with input for the same space
Overview
The context window is the fixed maximum number of tokens the model can attend over in one request — system prompt, retrieved documents, chat history, and the generated answer all share this budget. Exceed it and content must be truncated, summarized, or retrieved on demand.
How it works
- Start: Prompt Tokens System, user, retrieved context, and conversation history all consume tokens.
- Prompt Tokens -> Context Window The model can only attend over a fixed number of tokens at once.
- Context Window -> Budgeting Reserve room for the answer and rank context by relevance.
- Budgeting -> Truncate/Summarize Old or low-value tokens must be dropped, summarized, or retrieved on demand.
- Truncate/Summarize -> Output Tokens Long context helps recall, but increases latency, memory, and distraction risk.
In an interview
It's the hard cap on tokens per request, covering both input and output. Everything competes for one budget, so you rank context by relevance and reserve room for the completion. Bigger windows help recall but raise latency and cost, and quality can sag for facts buried in the middle of very long inputs.
Production defaults
- Allocation
- ~10% instructions, ~30% evidence, ~30% history, the rest free
- Ordering
- instructions first, evidence next, question last. The ends get the most attention
- Compact
- at ~70% full, summarize the oldest turns and drop the raw text
What breaks
- Longer prompts, worse answers — Dilution plus lost-in-the-middle. Cut and rerank rather than adding more.
- Output truncated mid-sentence — Input ate the window. Reserve output tokens explicitly and enforce it in code.