All concepts
Agentic RAG on AWS
Build an agentic RAG system on AWS with Bedrock Agents, Knowledge Bases, OpenSearch, and Guardrails.
Production AI Systems · Advanced · ~8 min
In plain English
The same RAG-plus-agent system you'd build by hand, assembled from AWS managed pieces: storage, embeddings, a vector index, a model endpoint and orchestration.
Why it's worth your time
It's the reference architecture you'll be asked to draw on a whiteboard, and knowing which managed piece maps to which concept is the whole answer.
If you remember three things
- Managed services buy speed and cost optionality
- The architecture is identical; only the boxes have vendor names
- Keep source documents and prompts outside the managed layer
Overview
A production agentic-RAG stack on AWS is fully serverless: API Gateway + Cognito front the request, a Lambda invokes a Bedrock Agent that decides when to retrieve, a Bedrock Knowledge Base handles chunking/embedding backed by OpenSearch Serverless for vector search, an Amazon Bedrock foundation model (e.g. Claude) generates the grounded answer, and Bedrock Guardrails filter it — all IAM-scoped and observable in CloudWatch.
How it works
- Authenticated request API Gateway with Cognito authenticates the caller and routes the request into the private VPC.
- Agent orchestrates A Lambda invokes a Bedrock Agent — a managed reasoning loop that decides when to retrieve and which action-group tools to call.
- Knowledge Base retrieves The agent queries a Bedrock Knowledge Base, which has already chunked and embedded your S3 documents.
- Vector search in OpenSearch The Knowledge Base runs kNN vector search over OpenSearch Serverless to fetch the most relevant passages.
- Grounded generation Retrieved passages are packed into the prompt and Claude on Bedrock generates an answer grounded in your data, with citations.
- Guardrails screen output Bedrock Guardrails redact PII, block denied topics, and run contextual-grounding checks to catch hallucination.
- Answer returns The grounded, filtered answer returns to the user — serverless, IAM-scoped, and traced end-to-end in CloudWatch.
In an interview
On AWS I front the API with API Gateway and Cognito, orchestrate with a Lambda calling a Bedrock Agent, retrieve through a Bedrock Knowledge Base backed by OpenSearch Serverless vector search, generate with a Bedrock model like Claude, and screen the output with Bedrock Guardrails. It's serverless, IAM-scoped, and monitored in CloudWatch — the managed-services path to agentic RAG.
Production defaults
- Shape
- object storage for documents → embedding + chunking pipeline → vector index → model endpoint → orchestration with tools
- Portability
- own the source corpus and the embedding pipeline. Re-indexing is a job; re-collecting a corpus is a project
- Cost
- embedding and re-indexing are the recurring line items people forget to model
- Security
- per-task IAM roles, VPC endpoints, and an audit trail on every action
What breaks
- Bill much higher than expected — Re-embedding on every document update. Embed on change only, and cache aggressively.
- Migration would take months — Your data lives only inside the managed index. Keep the source and the pipeline yourself.