All concepts

Vertex AI

Google Cloud's unified ML platform: training, serving, Vector Search, Gemini, and agents

Cloud Services · Intermediate · ~4 min

In plain English

Google's managed platform for training, tuning and serving models, including hosted access to foundation models.

Why it's worth your time

It's the GCP half of the 'draw the architecture' question, and the counterpart to SageMaker plus Bedrock combined.

If you remember three things

  • Training, tuning, endpoints and foundation models in one platform
  • Model Garden for pretrained and open-weights options
  • Endpoints bill for provisioned capacity

Overview

Vertex AI is Google Cloud's unified platform spanning both classic ML and generative AI. On the ML side it covers the lifecycle: notebooks, training on managed GPUs and TPUs, a Model Registry, autoscaling endpoints for online and batch prediction, and Model Monitoring that can retrigger training on drift. On the GenAI side it serves Gemini and other foundation models with a managed RAG Engine, ScaNN-powered Vector Search for billion-scale retrieval, and Agent Builder for assembling and deploying agents.

In an interview

Vertex AI unifies classic ML and GenAI on one platform, backed by data in Cloud Storage and BigQuery. You build in managed notebooks, train custom or AutoML models on managed GPUs/TPUs, govern them in the Model Registry, and serve on autoscaling endpoints, with Model Monitoring catching drift and skew to retrigger training. For GenAI it serves Gemini with a managed RAG Engine, Vector Search using Google's ScaNN for billion-scale retrieval, and Agent Builder to compose and deploy tool-using agents.

Production defaults

Serving
match endpoint type to traffic shape. Provisioned capacity for spiky traffic is wasted money
Tuning
adapter-based tuning before full fine-tuning, same logic as LoRA anywhere else
Registry
version models with their eval metrics attached
Portability
keep prompts, evals and source data in your own repo and storage

What breaks

  • Idle endpoint costs — Provisioned capacity bills whether or not you call it. Scale to zero or use batch prediction.
  • Quota errors at launch — GPU and model quotas are per-region and not generous by default. Request increases before launch day.

Watch it explained

What is Vertex AI? — Google Cloud Tech, 7:16

Related