All concepts

Context Engineering

Treat the window as a budget: decide what stays resident, compress the middle, and place what matters where the model still reads it.

Agentic Engineering · Intermediate · ~6 min

In plain English

Deciding what goes in the model's field of view this turn. Everything competes for the same space, and more is not better.

Why it's worth your time

Windows got huge, everyone stopped budgeting, and both quality and cost quietly got worse.

If you remember three things

  • Attention cost grows with the square of length
  • A fact buried mid-context is the one that gets missed
  • Compaction beats accumulation

Overview

Context engineering has quietly replaced prompt engineering as the job. The window is a budget that gets re-spent on every lap of the agent loop, so tokens are latency, cost and attention all at once. Four decisions make up the discipline. Window management: which blocks stay resident in the prompt versus fetched on demand. Compression: folding old turns into a summary that keeps the decisions and the numbers and throws away the prose. Context rot: understanding that recall collapses for facts buried mid-window, so a large window is not a large *usable* window. And ordering: scoring candidates on both recency and relevance, then placing the survivors where the model's attention actually lands.

How it works

  1. The window is a budget Every lap re-sends the whole prompt. Ask what earns its tokens on this lap, not how much fits.
  2. Resident vs retrieved Policy, current tools and recent turns stay. Corpora, old turns and big artefacts are fetched or written to files.
  3. Compress the middle Fold old turns into a running summary that preserves decisions and figures verbatim.
  4. Context rot Recall sags for facts buried mid-window. Measure needle-recall at your real length before trusting the spec sheet.
  5. Recency vs relevance Score on both, drop what earns nothing, then place survivors at the head and tail with the question last.

In an interview

Context engineering is treating the window as a budget that gets re-spent every lap. I split blocks into resident — system policy, the current task's tools, the last few turns — and retrieved, which is fetched only for the lap that needs it or offloaded to files. Old turns get compressed into a running summary that keeps decisions and figures verbatim and drops the prose. I assume context rot: recall sags for facts buried mid-window, so a 1M window is not 1M usable tokens and I measure needle recall at my real length. Then I score candidates on recency and relevance together, drop what earns nothing, and place the survivors at the head and tail with the question last.

Production defaults

Allocation
~10% instructions · ~30% evidence · ~30% history · rest free
Order
instructions first, evidence next, question last
Compact
at ~70% full, summarize old turns into a state note
Truncate tool output
hard. A 40k-token blob teaches nothing the first 2k didn't
Cache
stable prefix, volatile suffix, byte-identical, so prompt caching hits

What breaks

  • Quality falls as conversations lengthen — Dilution plus long-context degradation. Compact aggressively — shorter and relevant wins.
  • Prompt cache never hits — A timestamp or session id near the top changes the prefix every call. Move volatile content to the end.

Watch it explained

Context Engineering vs. Prompt Engineering: Smarter AI with RAG & Agents — IBM Technology, 7:52

Related