All concepts

Prompt Caching

The model already computed the KV for your system prompt — put the stable part first, mark it, and stop paying for it every call.

Advanced LLM Systems · Advanced · ~5 min

In plain English

You retell the same twenty-minute backstory before every conversation. Write it down once, hand it over, and start from the new part.

Why it's worth your time

Agent prompts are mostly identical every turn, so this is typically the largest single line on the bill and a big cut to time-to-first-token.

If you remember three things

  • Prefix-based: one changed token invalidates everything after it
  • Order stable → volatile, breakpoint after the stable part
  • Unlike semantic caching, it cannot change the answer

Overview

Every request re-runs the prefill over the entire prompt, and for an agent that prompt is mostly identical every turn: the same system instructions, the same tool schemas, the same retrieved document. Prompt caching stores the KV state for a prefix and reuses it, so a cache hit skips prefill for that span entirely — typically 90% cheaper on cached tokens and a large latency cut on long prompts. The rule that makes it work is ordering: caching is prefix-based, so anything that changes must come after everything that doesn't. Put a volatile timestamp at the top of a 20k-token system prompt and every request misses. The other rule is that cache entries expire on a short TTL, so caching only pays where traffic is steady.

In an interview

Prompt caching reuses the model's precomputed attention state for a prompt prefix, so repeated system instructions, tool schemas, and documents aren't re-processed on every call — usually around a tenth of the cost for those tokens and a big latency reduction. It's prefix-based, so the discipline is ordering: stable content first, volatile content last, with the cache breakpoint after the stable part.

Production defaults

Layout
system instructions → tool schemas → documents → [breakpoint] → conversation → user turn
Economics
cached input around 10% of normal, with a one-time write surcharge. Steady traffic pays, hourly traffic doesn't
Monitoring
alert on cache-read token ratio. A hit-rate regression is a cost incident no error rate will surface

What breaks

  • Cache hit rate is zero — Something volatile is at the top — a timestamp, a request id, or a reordered tool list. Anything above the breakpoint must be byte-identical.
  • Cost went up after enabling — Low-frequency traffic: entries expire before the next hit, so you pay the write surcharge every time. Measure before enabling.

Watch it explained

What is Prompt Caching? Optimize LLM Latency with AI Transformers — IBM Technology, 9:06

Related