All concepts

Zero-shot vs One-shot vs Few-shot

Control task behavior by giving no examples, one example, or several examples in the prompt.

Transformers & LLMs · Beginner · ~8 min

In plain English

Ask cold, show one example, or show several. Each extra example costs context but teaches format and edge cases faster than any instruction.

Why it's worth your time

It's usually the first thing to try and the last thing people optimize — a well-chosen example beats a paragraph of instructions.

If you remember three things

  • Examples teach FORMAT more reliably than they teach reasoning
  • 3–5 examples captures most of the benefit
  • Example order matters, and the last one matters most

Overview

Three points on the in-context learning dial. Zero-shot gives only a task instruction, one-shot adds a single input→output demonstration, and few-shot supplies several. The model infers the desired mapping from whatever sits in the prompt — no weight updates, just pattern-matching over context.

How it works

  1. Start: Instruction Zero-shot uses only a task instruction.
  2. Instruction -> One Example One-shot shows a single input-output pattern.
  3. One Example -> Few Examples Few-shot adds multiple demonstrations to clarify format and edge cases.
  4. Few Examples -> Pattern Match The model infers the desired mapping from examples in context.
  5. Pattern Match -> Answer Examples improve reliability but consume context and can bias the response.

In an interview

It's how many worked examples you put in the prompt: none, one, or a handful. More demonstrations pin down the output format and edge cases, raising reliability — but each example burns context tokens and can bias the model toward superficial patterns in your samples.

Production defaults

Count
start zero-shot; add examples only where it actually fails. 3–5 is the usual sweet spot
Selection
cover the edge cases, not the easy middle. Dynamic selection by similarity to the query beats a fixed set
Caching
put fixed examples in the stable prefix so prompt caching covers them

What breaks

  • The model copies the examples too literally — Examples too similar to each other. Vary them, or say explicitly that they illustrate format only.
  • Adding more examples made it worse — Context dilution, and a bias toward the last example. Fewer, better-chosen examples.

Watch it explained

Mastering Prompting Techniques: Zero Shot, Few Shot, & Chain of Thought (COT) Explained — The Coding Neuron, 6:23

Related