All concepts

Constrained Decoding

Mask the logits so only tokens that keep the output valid can be sampled — schema conformance by construction, not by retrying.

Advanced LLM Systems · Advanced · ~6 min

In plain English

Writing on a form where the boxes only accept the right kind of character. You can't put letters in a date field, so you never have to check afterwards.

Why it's worth your time

It deletes the parse-retry loop entirely: malformed output stops being unlikely and starts being impossible.

If you remember three things

  • Schema → regex → FSM → per-state allowed-token mask
  • The model still chooses; it just can't choose invalid
  • Guarantees syntax, never semantics

Overview

Asking politely for JSON and retrying on parse errors is a probability game you lose at scale. Constrained decoding makes invalid output unrepresentable. Compile the schema or grammar into a finite-state machine, and at each decoding step compute the set of tokens that can legally follow the current state; set every other logit to negative infinity before sampling. The model still chooses freely among valid continuations, so its judgment is preserved — it simply cannot emit a stray comma or an unquoted key. Libraries like Outlines precompute the token-mask index so the per-step overhead is a lookup, and llama.cpp's GBNF and vLLM's guided decoding expose the same idea. The residual risk is semantic, not syntactic: valid JSON can still be wrong.

In an interview

Constrained decoding compiles your schema or grammar into a state machine and, at every token, masks out any token that would make the output invalid. The model still picks among the legal options, so quality is preserved, but malformed JSON becomes structurally impossible rather than merely unlikely. It removes the parse-retry loop entirely — though it guarantees syntax, not semantics.

Production defaults

Schema shape
put a free-text `reasoning` field FIRST, constrained fields after. Over-constraining a reasoning task measurably hurts answers
Tooling
Outlines / XGrammar locally, or the provider's native structured-output mode when hosted
Validation
still validate meaning after parsing — enum values can be schema-valid and wrong

What breaks

  • Answers got worse after adding a schema — The model has nowhere to think. Add a reasoning field before the structured fields.
  • Valid JSON, wrong values — Working as designed — the FSM only enforces shape. Semantic validation belongs downstream.

Watch it explained

Constrained Decoding Explained: How LLMs Generate Perfect Structured Output — Engineering Insider, 8:51

Related