All concepts
Model Routing & Token Cost
Send easy queries to cheap models and hard ones to strong models to cut cost without losing quality.
MLOps & LLMOps · Advanced · ~8 min
In plain English
Send easy requests to a small cheap model and hard ones to a big one, so you're not paying premium prices for 'what time do you open?'
Why it's worth your time
It routinely takes 50–70% off a model bill with no visible quality change — the highest-leverage cost lever there is.
If you remember three things
- Classify difficulty, then route
- Escalate on low confidence rather than guessing up front
- Log every routing decision or you can never tune it
Overview
Not every request needs your most expensive model. A router classifies query difficulty (or tries a cheap model first) and escalates only when needed. Combined with caching and prompt trimming, routing is the biggest lever on LLM cost and latency.
How it works
- Check the cache first A semantic cache lookup: if we've answered something similar, return it instantly — zero tokens.
- Classify difficulty A cheap classifier predicts how hard the query is — the routing decision.
- Easy → cheap model Simple, common queries go to a small/cheap model. Most traffic is easy, so this is where the savings come from.
- Hard → frontier model Only genuinely complex queries escalate to the strong (expensive) model.
- Verify the hard path A quality check catches the risk case — a hard query a cheaper tier would have failed.
- Return the answer Routing plus caching cuts cost and latency substantially with quality held flat.
In an interview
Model routing sends each request to the cheapest model that can handle it — a small/cheap model for easy queries, a frontier model for hard ones — using a classifier or a cheap-first cascade with a quality check. With caching and prompt trimming, it's the main way to cut LLM cost and latency without hurting quality.
Production defaults
- Classifier
- a cheap model or a heuristic. Don't spend a big-model call deciding which model to use
- Escalation
- small model first, escalate on low confidence or an explicit 'I'm not sure'
- Track
- escalation rate as a dashboard metric, and sample escalations to check they needed it
- Fallback
- always have one for provider outages. Routing infrastructure gives you this for free
What breaks
- Everything routes to the big model — Confidence threshold tuned for safety and never revisited. Sample escalations and retune.
- Quality dropped on a subset of traffic — A category that looked easy isn't. Route by task type, not just by length.