All concepts

Context Window

The maximum tokens the model can read and generate in one request.

Transformers & LLMs · Beginner · ~8 min

In plain English

The model's desk. Everything it can see this turn has to fit on it — instructions, history, retrieved documents, and room to write the answer.

Why it's worth your time

Huge windows made people stop budgeting, and both quality and cost quietly got worse. The window is a budget.

If you remember three things

  • Attention cost grows with the square of the length
  • Facts in the middle of a long context are the ones missed
  • Output tokens compete with input for the same space

Overview

The context window is the fixed maximum number of tokens the model can attend over in one request — system prompt, retrieved documents, chat history, and the generated answer all share this budget. Exceed it and content must be truncated, summarized, or retrieved on demand.

How it works

  1. Start: Prompt Tokens System, user, retrieved context, and conversation history all consume tokens.
  2. Prompt Tokens -> Context Window The model can only attend over a fixed number of tokens at once.
  3. Context Window -> Budgeting Reserve room for the answer and rank context by relevance.
  4. Budgeting -> Truncate/Summarize Old or low-value tokens must be dropped, summarized, or retrieved on demand.
  5. Truncate/Summarize -> Output Tokens Long context helps recall, but increases latency, memory, and distraction risk.

In an interview

It's the hard cap on tokens per request, covering both input and output. Everything competes for one budget, so you rank context by relevance and reserve room for the completion. Bigger windows help recall but raise latency and cost, and quality can sag for facts buried in the middle of very long inputs.

Production defaults

Allocation
~10% instructions, ~30% evidence, ~30% history, the rest free
Ordering
instructions first, evidence next, question last. The ends get the most attention
Compact
at ~70% full, summarize the oldest turns and drop the raw text

What breaks

  • Longer prompts, worse answers — Dilution plus lost-in-the-middle. Cut and rerank rather than adding more.
  • Output truncated mid-sentence — Input ate the window. Reserve output tokens explicitly and enforce it in code.

Watch it explained

Tokens Explained in AI & ChatGPT | Context Window, Tokenization & LLM Tokens — Code with Neema, 5:03

Related