All concepts

Agentic RAG on AWS

Build an agentic RAG system on AWS with Bedrock Agents, Knowledge Bases, OpenSearch, and Guardrails.

Production AI Systems · Advanced · ~8 min

In plain English

The same RAG-plus-agent system you'd build by hand, assembled from AWS managed pieces: storage, embeddings, a vector index, a model endpoint and orchestration.

Why it's worth your time

It's the reference architecture you'll be asked to draw on a whiteboard, and knowing which managed piece maps to which concept is the whole answer.

If you remember three things

  • Managed services buy speed and cost optionality
  • The architecture is identical; only the boxes have vendor names
  • Keep source documents and prompts outside the managed layer

Overview

A production agentic-RAG stack on AWS is fully serverless: API Gateway + Cognito front the request, a Lambda invokes a Bedrock Agent that decides when to retrieve, a Bedrock Knowledge Base handles chunking/embedding backed by OpenSearch Serverless for vector search, an Amazon Bedrock foundation model (e.g. Claude) generates the grounded answer, and Bedrock Guardrails filter it — all IAM-scoped and observable in CloudWatch.

How it works

  1. Authenticated request API Gateway with Cognito authenticates the caller and routes the request into the private VPC.
  2. Agent orchestrates A Lambda invokes a Bedrock Agent — a managed reasoning loop that decides when to retrieve and which action-group tools to call.
  3. Knowledge Base retrieves The agent queries a Bedrock Knowledge Base, which has already chunked and embedded your S3 documents.
  4. Vector search in OpenSearch The Knowledge Base runs kNN vector search over OpenSearch Serverless to fetch the most relevant passages.
  5. Grounded generation Retrieved passages are packed into the prompt and Claude on Bedrock generates an answer grounded in your data, with citations.
  6. Guardrails screen output Bedrock Guardrails redact PII, block denied topics, and run contextual-grounding checks to catch hallucination.
  7. Answer returns The grounded, filtered answer returns to the user — serverless, IAM-scoped, and traced end-to-end in CloudWatch.

In an interview

On AWS I front the API with API Gateway and Cognito, orchestrate with a Lambda calling a Bedrock Agent, retrieve through a Bedrock Knowledge Base backed by OpenSearch Serverless vector search, generate with a Bedrock model like Claude, and screen the output with Bedrock Guardrails. It's serverless, IAM-scoped, and monitored in CloudWatch — the managed-services path to agentic RAG.

Production defaults

Shape
object storage for documents → embedding + chunking pipeline → vector index → model endpoint → orchestration with tools
Portability
own the source corpus and the embedding pipeline. Re-indexing is a job; re-collecting a corpus is a project
Cost
embedding and re-indexing are the recurring line items people forget to model
Security
per-task IAM roles, VPC endpoints, and an audit trail on every action

What breaks

  • Bill much higher than expected — Re-embedding on every document update. Embed on change only, and cache aggressively.
  • Migration would take months — Your data lives only inside the managed index. Keep the source and the pipeline yourself.

Watch it explained

Knowledge Bases for Amazon Bedrock: Chat with your Document — AWS Developers, 4:11

Related