Treat the window as a budget: decide what stays resident, compress the middle, and place what matters where the model still reads it.
Deciding what goes in the model's field of view this turn. Everything competes for the same space, and more is not better.
Windows got huge, everyone stopped budgeting, and both quality and cost quietly got worse.
Context engineering has quietly replaced prompt engineering as the job. The window is a budget that gets re-spent on every lap of the agent loop, so tokens are latency, cost and attention all at once. Four decisions make up the discipline. Window management: which blocks stay resident in the prompt versus fetched on demand. Compression: folding old turns into a summary that keeps the decisions and the numbers and throws away the prose. Context rot: understanding that recall collapses for facts buried mid-window, so a large window is not a large *usable* window. And ordering: scoring candidates on both recency and relevance, then placing the survivors where the model's attention actually lands.
Context engineering is treating the window as a budget that gets re-spent every lap. I split blocks into resident — system policy, the current task's tools, the last few turns — and retrieved, which is fetched only for the lap that needs it or offloaded to files. Old turns get compressed into a running summary that keeps decisions and figures verbatim and drops the prose. I assume context rot: recall sags for facts buried mid-window, so a 1M window is not 1M usable tokens and I measure needle recall at my real length. Then I score candidates on recency and relevance together, drop what earns nothing, and place the survivors at the head and tail with the question last.
Context Engineering vs. Prompt Engineering: Smarter AI with RAG & Agents — IBM Technology, 7:52