Mask the logits so only tokens that keep the output valid can be sampled — schema conformance by construction, not by retrying.
Writing on a form where the boxes only accept the right kind of character. You can't put letters in a date field, so you never have to check afterwards.
It deletes the parse-retry loop entirely: malformed output stops being unlikely and starts being impossible.
Asking politely for JSON and retrying on parse errors is a probability game you lose at scale. Constrained decoding makes invalid output unrepresentable. Compile the schema or grammar into a finite-state machine, and at each decoding step compute the set of tokens that can legally follow the current state; set every other logit to negative infinity before sampling. The model still chooses freely among valid continuations, so its judgment is preserved — it simply cannot emit a stray comma or an unquoted key. Libraries like Outlines precompute the token-mask index so the per-step overhead is a lookup, and llama.cpp's GBNF and vLLM's guided decoding expose the same idea. The residual risk is semantic, not syntactic: valid JSON can still be wrong.
Constrained decoding compiles your schema or grammar into a state machine and, at every token, masks out any token that would make the output invalid. The model still picks among the legal options, so quality is preserved, but malformed JSON becomes structurally impossible rather than merely unlikely. It removes the parse-retry loop entirely — though it guarantees syntax, not semantics.
Constrained Decoding Explained: How LLMs Generate Perfect Structured Output — Engineering Insider, 8:51