All concepts
Semantic Cache
Reuse prior answers or retrieved contexts when a new query is semantically similar.
RAG & Retrieval · Advanced · ~8 min
In plain English
Remember answers to questions you've already been asked — and recognize a question you've seen before even when it's worded differently.
Why it's worth your time
On real traffic a large share of questions repeat, and a cache hit is free and instant.
If you remember three things
- Exact-match cache first; semantic match on top
- A wrong cache hit is worse than a miss
- Never cache across users or permission boundaries
Overview
A cache keyed by meaning, not exact string. Incoming queries are embedded and matched against stored requests; a close-enough hit with matching metadata reuses the prior answer or retrieved context, skipping the LLM call. Misses run the full pipeline and write back the result with TTL and policy tags.
How it works
- Start: New Query The user asks something that may be close to an earlier request.
- New Query -> Query Embedding Embed the query and search a cache index.
- Query Embedding -> Similarity Hit If distance is above threshold and metadata matches, reuse cached answer/context.
- Similarity Hit -> Miss -> RAG If not safe, run the full retrieval and generation pipeline.
- Miss -> RAG -> Write Cache Store answer, sources, model version, TTL, tenant, and policy metadata.
In an interview
A semantic cache reuses past answers when a new query is close in embedding space, not just identical text. You embed the query, search the cache, and if similarity clears a threshold and metadata matches, return the cached result — cutting latency and cost. The main risk is false hits from a loose threshold.
Production defaults
- Threshold
- cosine ~0.95. High on purpose — false hits are expensive
- Key scope
- include tenant, user permissions, and model version in the cache key
- TTL
- short for anything time-sensitive. Stale correct answers are still wrong answers
What breaks
- A user saw another user's answer — Cache key missing the tenant/permission scope. This is a security bug, not a caching bug.
- Near-miss questions get the wrong cached answer — Threshold too low. 'How do I cancel?' and 'How do I cancel my refund?' are close vectors, different questions.