The model already computed the KV for your system prompt — put the stable part first, mark it, and stop paying for it every call.
You retell the same twenty-minute backstory before every conversation. Write it down once, hand it over, and start from the new part.
Agent prompts are mostly identical every turn, so this is typically the largest single line on the bill and a big cut to time-to-first-token.
Every request re-runs the prefill over the entire prompt, and for an agent that prompt is mostly identical every turn: the same system instructions, the same tool schemas, the same retrieved document. Prompt caching stores the KV state for a prefix and reuses it, so a cache hit skips prefill for that span entirely — typically 90% cheaper on cached tokens and a large latency cut on long prompts. The rule that makes it work is ordering: caching is prefix-based, so anything that changes must come after everything that doesn't. Put a volatile timestamp at the top of a 20k-token system prompt and every request misses. The other rule is that cache entries expire on a short TTL, so caching only pays where traffic is steady.
Prompt caching reuses the model's precomputed attention state for a prompt prefix, so repeated system instructions, tool schemas, and documents aren't re-processed on every call — usually around a tenth of the cost for those tokens and a big latency reduction. It's prefix-based, so the discipline is ordering: stable content first, volatile content last, with the cache breakpoint after the stable part.
What is Prompt Caching? Optimize LLM Latency with AI Transformers — IBM Technology, 9:06