Buy accuracy with inference tokens instead of training runs — think longer, sample more, verify, and pick.
Sitting an exam. You can answer straight off, or take the extra twenty minutes to check your working. On hard questions the extra time is worth more than a smarter pen.
It is the second scaling axis: you can buy accuracy at inference without retraining anything — but only on problems that actually need reasoning.
The second scaling axis is inference. A model that spends more tokens before answering — extended thinking, sampling several candidates, verifying and selecting — measurably outperforms itself at a single greedy pass, and on hard reasoning tasks the gain can exceed what a much larger model gives you at fixed latency. The three levers are sequential (longer chains, self-revision), parallel (best-of-n with a verifier or majority vote), and search (branch, evaluate, prune). Returns are strongly diminishing and task-dependent: doubling the thinking budget on a lookup question buys nothing but latency. The engineering problem is therefore allocation — spend the budget where difficulty warrants it, which is what makes an explicit thinking budget and a difficulty router part of the design.
Test-time compute is the idea that you can trade inference tokens for accuracy. Sequentially, the model thinks longer or revises itself; in parallel, you sample several answers and let a verifier or majority vote choose. On hard reasoning tasks this can beat using a much larger model at the same latency, but returns diminish sharply and easy questions gain nothing — so the actual engineering is deciding which requests deserve the budget.
What Are Large Reasoning Models (LRMs)? Smarter AI Beyond LLMs — IBM Technology, 8:38