All concepts
Agentic RAG on GCP
Build an agentic RAG system on GCP with Cloud Run, Vertex AI Agent Builder, RAG Engine, Vector Search, and Gemini.
Production AI Systems · Advanced · ~8 min
In plain English
The same architecture again on Google Cloud: storage, an embedding and index service, a model endpoint, and the orchestration around them.
Why it's worth your time
Seeing the identical architecture in two clouds is what makes the pattern portable in your head rather than vendor-specific.
If you remember three things
- Same five boxes, different service names
- BigQuery makes analytics over traces and feedback easy
- The lock-in risk is the index, not the model
Overview
A production agentic-RAG stack on Google Cloud runs the agent on Cloud Run (scales to zero) behind a Global Load Balancer, orchestrated with Vertex AI Agent Builder / the Agent Development Kit. Retrieval uses the Vertex AI RAG Engine backed by Vertex AI Vector Search (ScaNN), generation uses Gemini on Vertex with grounding + citations, and Vertex safety filters / Model Armor screen the output — observable in Cloud Logging and Trace.
How it works
- Request routed A Global Load Balancer terminates TLS and routes the request to a Cloud Run service.
- Agent on Cloud Run Cloud Run hosts the agent (Vertex AI Agent Builder / ADK) and scales to zero when idle — you pay only for traffic.
- RAG Engine retrieves The agent calls the Vertex AI RAG Engine, which manages corpora, chunking, and embeddings for you.
- Vector Search (ScaNN) RAG Engine retrieves candidates from Vertex AI Vector Search — Google's ScaNN ANN index — for fast semantic search.
- Gemini grounds the answer Context is grounded into Gemini on Vertex, which generates the answer with grounding metadata and citations.
- Safety screening Vertex AI safety filters and Model Armor screen the output for content safety and prompt-injection.
- Answer returns The cited, grounded answer returns — serverless on Cloud Run, IAM/VPC-SC scoped, and traced in Cloud Logging.
In an interview
On GCP I run the agent on Cloud Run behind a Global Load Balancer, orchestrate with Vertex AI Agent Builder, retrieve via the Vertex RAG Engine backed by Vertex Vector Search (ScaNN), generate with Gemini on Vertex using grounding, and screen output with Vertex safety filters / Model Armor. Serverless, autoscaling to zero, and traced in Cloud Logging — the managed-services path to agentic RAG on Google Cloud.
Production defaults
- Shape
- Cloud Storage → embedding + chunking → vector index → model endpoint → orchestration
- Analytics
- land traces and feedback in a warehouse from day one. The feedback loop needs somewhere to look
- Portability
- keep documents, prompts and evals outside the managed services
- Security
- per-task service accounts, least privilege, full audit logging
What breaks
- Two clouds, two architectures — They shouldn't be. If your design doesn't map across, the vendor is in your architecture.
- Index rebuild takes days — Plan for re-indexing as a routine operation — embedding model changes force it.