All concepts

RAG Pipeline

Retrieve relevant chunks for a query, stuff them into the prompt, and let the LLM answer with citations.

RAG & Retrieval · Intermediate · ~11 min

In plain English

Before answering, go and look it up. Fetch the most relevant passages from your own documents and hand them to the model as the material to answer from.

Why it's worth your time

It's the cheapest way to give a model private, current, citable knowledge — and the system design question you will be asked in every interview.

If you remember three things

  • Index ahead of time; retrieve and answer at query time
  • Retrieval failures, not generation failures, dominate RAG errors
  • Citations are what make the answer auditable

Overview

Retrieval-Augmented Generation grounds an LLM in your data. Documents are chunked and embedded into a vector store ahead of time. At query time you embed the question, retrieve the most similar chunks, optionally rerank them, and feed them as context so the model answers from real sources instead of memory — reducing hallucination and enabling citations.

How it works

  1. User query A question arrives. The LLM alone might hallucinate or lack your private/up-to-date data.
  2. Embed the query Convert the query into a vector using the same embedding model used to index the documents.
  3. Retrieve top-k Find the k nearest chunks in the vector store by cosine similarity (often via an ANN index like HNSW).
  4. Rerank A cross-encoder rescoring the query against each chunk sharpens ordering, pushing the truly relevant chunks up.
  5. Augment the prompt Insert the retrieved chunks into the prompt as context, with instructions to answer only from them and cite sources.
  6. Generate with citations The LLM produces a grounded answer that cites the chunks it used — traceable and less prone to hallucination.

In an interview

RAG grounds an LLM in external data: chunk and embed documents into a vector store, then at query time embed the question, retrieve the most similar chunks, optionally rerank, and feed them into the prompt so the model answers from sources with citations. It reduces hallucination and handles private or fresh data without retraining.

Production defaults

Chunks
600–1000 tokens, ~15% overlap, split on structure
Retrieve
top-50 hybrid (BM25 + vector) → rerank → top-3 to 5
Context
keep evidence under ~30% of the window
Prompt
'Answer only from the context; if it isn't there, say you don't know' + required citations

What breaks

  • Confident but wrong answers — The passage was never retrieved. Measure retrieval hit-rate separately — no prompt fixes a retrieval miss.
  • Ignores a fact that IS in the context — Lost in the middle. Rerank so the best passage is first, and cut k rather than raising it.

Watch it explained

What is Retrieval-Augmented Generation (RAG)? — IBM Technology, 6:35

Related