All concepts

LLM as a Judge

Use a stronger model with a rubric to grade answers at scale.

Model Evaluation · Advanced · ~8 min

In plain English

Use a strong model to grade another model's answers against a written rubric — the way you'd hand a marking scheme to a teaching assistant.

Why it's worth your time

It's the only practical way to score open-ended output at volume, and the backbone of every serious LLM eval pipeline.

If you remember three things

  • A rubric with anchored levels beats 'rate 1–10'
  • Ask for the reason before the score
  • Calibrate against human labels or the number means nothing

Overview

Using a capable model, guided by an explicit rubric, to score other models' outputs at scale. The judge grades dimensions like relevance, factuality, and format, replacing slow human review — but it must be calibrated against human labels to be trusted.

How it works

  1. Start: Eval Prompt Define the rubric, task, reference answer, and exact scoring dimensions.
  2. Eval Prompt -> Candidate Answer Feed the model output plus context into the judge.
  3. Candidate Answer -> Judge Model A capable LLM scores relevance, factuality, format, safety, or helpfulness.
  4. Judge Model -> Calibration Compare judge scores to human labels and use pairwise or blinded judging to reduce bias.
  5. Calibration -> Eval Metric Aggregate scores into release gates, regression tests, and production monitoring.

In an interview

An evaluation method where a strong LLM grades candidate answers against a rubric and optional reference, scoring factuality, helpfulness, or safety. It scales far past human labeling, but the judge carries biases (position, verbosity, self-preference), so I use pairwise or blinded comparisons and check agreement with human labels.

Production defaults

Rubric
explicit criteria, a defined scale, one dimension at a time
Order
reason first, then score. Score-first answers are post-hoc rationalized
Calibration
hand-label ~50 examples; below ~80% agreement, fix the rubric before trusting the judge
Bias control
randomize position in pairwise comparisons, hide provenance, judge with a different model family than the one under test

What breaks

  • Everything scores 4/5 — Vague rubric. Anchor each level with a concrete example of what earns it.
  • The judge prefers its own family's answers — Documented self-preference bias. Swap judge family and randomize presentation order.

Watch it explained

LLM as a Judge: Scaling AI Evaluation Strategies — IBM Technology, 6:08

Related