All concepts

Model Routing & Token Cost

Send easy queries to cheap models and hard ones to strong models to cut cost without losing quality.

MLOps & LLMOps · Advanced · ~8 min

In plain English

Send easy requests to a small cheap model and hard ones to a big one, so you're not paying premium prices for 'what time do you open?'

Why it's worth your time

It routinely takes 50–70% off a model bill with no visible quality change — the highest-leverage cost lever there is.

If you remember three things

  • Classify difficulty, then route
  • Escalate on low confidence rather than guessing up front
  • Log every routing decision or you can never tune it

Overview

Not every request needs your most expensive model. A router classifies query difficulty (or tries a cheap model first) and escalates only when needed. Combined with caching and prompt trimming, routing is the biggest lever on LLM cost and latency.

How it works

  1. Check the cache first A semantic cache lookup: if we've answered something similar, return it instantly — zero tokens.
  2. Classify difficulty A cheap classifier predicts how hard the query is — the routing decision.
  3. Easy → cheap model Simple, common queries go to a small/cheap model. Most traffic is easy, so this is where the savings come from.
  4. Hard → frontier model Only genuinely complex queries escalate to the strong (expensive) model.
  5. Verify the hard path A quality check catches the risk case — a hard query a cheaper tier would have failed.
  6. Return the answer Routing plus caching cuts cost and latency substantially with quality held flat.

In an interview

Model routing sends each request to the cheapest model that can handle it — a small/cheap model for easy queries, a frontier model for hard ones — using a classifier or a cheap-first cascade with a quality check. With caching and prompt trimming, it's the main way to cut LLM cost and latency without hurting quality.

Production defaults

Classifier
a cheap model or a heuristic. Don't spend a big-model call deciding which model to use
Escalation
small model first, escalate on low confidence or an explicit 'I'm not sure'
Track
escalation rate as a dashboard metric, and sample escalations to check they needed it
Fallback
always have one for provider outages. Routing infrastructure gives you this for free

What breaks

  • Everything routes to the big model — Confidence threshold tuned for safety and never revisited. Sample escalations and retune.
  • Quality dropped on a subset of traffic — A category that looked easy isn't. Route by task type, not just by length.

Watch it explained

What is an LLM Router? — Sam Witteveen, 9:16

Related