All concepts
LLM as a Judge
Use a stronger model with a rubric to grade answers at scale.
Model Evaluation · Advanced · ~8 min
In plain English
Use a strong model to grade another model's answers against a written rubric — the way you'd hand a marking scheme to a teaching assistant.
Why it's worth your time
It's the only practical way to score open-ended output at volume, and the backbone of every serious LLM eval pipeline.
If you remember three things
- A rubric with anchored levels beats 'rate 1–10'
- Ask for the reason before the score
- Calibrate against human labels or the number means nothing
Overview
Using a capable model, guided by an explicit rubric, to score other models' outputs at scale. The judge grades dimensions like relevance, factuality, and format, replacing slow human review — but it must be calibrated against human labels to be trusted.
How it works
- Start: Eval Prompt Define the rubric, task, reference answer, and exact scoring dimensions.
- Eval Prompt -> Candidate Answer Feed the model output plus context into the judge.
- Candidate Answer -> Judge Model A capable LLM scores relevance, factuality, format, safety, or helpfulness.
- Judge Model -> Calibration Compare judge scores to human labels and use pairwise or blinded judging to reduce bias.
- Calibration -> Eval Metric Aggregate scores into release gates, regression tests, and production monitoring.
In an interview
An evaluation method where a strong LLM grades candidate answers against a rubric and optional reference, scoring factuality, helpfulness, or safety. It scales far past human labeling, but the judge carries biases (position, verbosity, self-preference), so I use pairwise or blinded comparisons and check agreement with human labels.
Production defaults
- Rubric
- explicit criteria, a defined scale, one dimension at a time
- Order
- reason first, then score. Score-first answers are post-hoc rationalized
- Calibration
- hand-label ~50 examples; below ~80% agreement, fix the rubric before trusting the judge
- Bias control
- randomize position in pairwise comparisons, hide provenance, judge with a different model family than the one under test
What breaks
- Everything scores 4/5 — Vague rubric. Anchor each level with a concrete example of what earns it.
- The judge prefers its own family's answers — Documented self-preference bias. Swap judge family and randomize presentation order.