All concepts
Chain-of-Thought & Self-Consistency
Make the model reason step by step, then vote across several chains for a reliable answer.
Transformers & LLMs · Intermediate · ~8 min
In plain English
Ask the model to show its working. Because it only has one forward pass per token, writing the intermediate steps is literally how it gets more thinking done.
Why it's worth your time
It's the highest-leverage prompting change available, and understanding WHY it works stops you using it where it doesn't.
If you remember three things
- The reasoning tokens are the computation, not a report of it
- Biggest gains on multi-step arithmetic and logic
- The written reasoning is not guaranteed to be the real cause of the answer
Overview
Chain-of-thought (CoT) prompts the model to produce intermediate reasoning before the final answer, which sharply improves multi-step tasks. Self-consistency samples several independent chains at higher temperature and takes the majority-vote answer — correct reasoning paths tend to agree, wrong ones scatter.
How it works
- Ask for reasoning Chain-of-thought prompts the model to reason step by step before committing to an answer.
- Sample several chains Self-consistency samples multiple independent reasoning chains at a higher temperature.
- Each reaches an answer Different chains take different routes but often converge on the same final answer.
- Majority vote Take the most common final answer — wrong reasoning paths disagree, right ones agree.
- Final answer Self-consistency reliably beats a single greedy chain on math and logic.
In an interview
Chain-of-thought asks the model to think step by step before answering, which helps a lot on math and logic. Self-consistency improves it further: sample multiple reasoning chains and take the majority answer. Right reasoning tends to converge on the same answer, wrong reasoning disagrees, so voting beats a single greedy chain.
Production defaults
- Prompt
- 'Think step by step, then give the final answer after ---'. Parse only what follows the marker
- Self-consistency
- sample several chains and take the majority answer when accuracy matters more than cost
- Reasoning models
- already do this internally. Adding 'think step by step' to them is redundant and can hurt
What breaks
- Reasoning looks right, answer is wrong — Parse the final answer explicitly with a delimiter — don't regex the last number out of the prose.
- Latency and cost tripled — You're paying for every reasoning token. Use CoT selectively, on the hard slice only.