All concepts

Test-Time Compute & Reasoning Models

Buy accuracy with inference tokens instead of training runs — think longer, sample more, verify, and pick.

Advanced LLM Systems · Advanced · ~7 min

In plain English

Sitting an exam. You can answer straight off, or take the extra twenty minutes to check your working. On hard questions the extra time is worth more than a smarter pen.

Why it's worth your time

It is the second scaling axis: you can buy accuracy at inference without retraining anything — but only on problems that actually need reasoning.

If you remember three things

  • Sequential (think longer) vs parallel (best-of-n) vs search
  • Accuracy grows roughly with log(compute) — returns flatten fast
  • Best-of-n can never be better than its verifier

Overview

The second scaling axis is inference. A model that spends more tokens before answering — extended thinking, sampling several candidates, verifying and selecting — measurably outperforms itself at a single greedy pass, and on hard reasoning tasks the gain can exceed what a much larger model gives you at fixed latency. The three levers are sequential (longer chains, self-revision), parallel (best-of-n with a verifier or majority vote), and search (branch, evaluate, prune). Returns are strongly diminishing and task-dependent: doubling the thinking budget on a lookup question buys nothing but latency. The engineering problem is therefore allocation — spend the budget where difficulty warrants it, which is what makes an explicit thinking budget and a difficulty router part of the design.

In an interview

Test-time compute is the idea that you can trade inference tokens for accuracy. Sequentially, the model thinks longer or revises itself; in parallel, you sample several answers and let a verifier or majority vote choose. On hard reasoning tasks this can beat using a much larger model at the same latency, but returns diminish sharply and easy questions gain nothing — so the actual engineering is deciding which requests deserve the budget.

Production defaults

Budgets
per request class: none for lookups, 2-4k thinking tokens for analysis, more only where an eval shows it pays
Best-of-n
n = 5-8 with a real verifier. Majority vote only when the answer is checkable
Cost
budget at p95, not the mean — variable-compute inference makes averages misleading

What breaks

  • Cost tripled, quality flat — Thinking is on globally. Route by difficulty; easy traffic gains nothing from a reasoning budget.
  • best-of-n picks bad answers — Your verifier is weaker than your generator. Selection quality is the ceiling — fix the verifier or drop the technique.

Watch it explained

What Are Large Reasoning Models (LRMs)? Smarter AI Beyond LLMs — IBM Technology, 8:38

Related