All concepts
Citations & Grounding
Tie generated claims back to retrieved source chunks so answers are auditable.
RAG & Retrieval · Intermediate · ~8 min
In plain English
Make the model point at the exact sentence it used. If it can't point, it probably made it up.
Why it's worth your time
It's the difference between an answer a user believes and an answer a user can check — and it's usually a compliance requirement, not a nicety.
If you remember three things
- Cite spans, not whole documents
- Verify the citation actually supports the claim, in code
- 'I don't know' must be an allowed, rewarded answer
Overview
Grounding ties every generated claim back to the retrieved chunk that supports it. Each source carries a stable ID (URL, page, row, or span); the system maps claims to their smallest supporting evidence and surfaces clickable citations, making answers auditable and hallucinations detectable.
How it works
- Start: Answer Claims The LLM produces claims that must be supported by context.
- Answer Claims -> Retrieved Chunks Each chunk has a stable source ID, URL, page, row, or document span.
- Retrieved Chunks -> Claim-Source Map Claims are linked to the smallest supporting chunks.
- Claim-Source Map -> Citations The UI renders citations next to claims and can open the original source.
- Citations -> Audit Trail Grounding metadata is logged for evaluation and debugging.
In an interview
It's the discipline of linking each claim in an answer to the exact retrieved chunk that backs it. You attach stable source IDs to chunks, map claims to the smallest supporting span, and render citations in the UI — so users can verify and you can log an audit trail.
Production defaults
- Format
- give each retrieved chunk an id and require inline references to those ids
- Verify
- post-check that each cited id was actually retrieved. Fabricated citations are a real failure mode
- Refusal
- explicitly permit 'not in the provided context'. Without it, the model will invent rather than decline
- Measure
- citation precision (does it support the claim) separately from answer quality
What breaks
- Citations point at the wrong passage — The model is attaching plausible ids after the fact. Verify id-to-claim support programmatically and reject on failure.
- The model never says 'I don't know' — Nothing rewards it. Say so explicitly in the prompt, and include unanswerable cases in your eval set.