All concepts

AWS SageMaker

AWS's managed ML platform: train on GPUs, version models, and serve them as endpoints

Cloud Services · Intermediate · ~4 min

In plain English

Managed infrastructure for the whole model lifecycle — training jobs, tuning, and endpoints — so you're not building GPU orchestration yourself.

Why it's worth your time

It's the default answer to 'how do we train and serve this on AWS' and shows up in every AWS-flavoured system design round.

If you remember three things

  • Training jobs, model registry, and endpoints are separate pieces
  • Endpoints bill for uptime, not per request
  • Bring your own container when the managed image doesn't fit

Overview

SageMaker is Amazon's managed platform for the whole ML lifecycle. You build in Studio notebooks, launch training jobs that spin up a managed GPU cluster on demand and tear it down when done, version artifacts in the Model Registry, and deploy either as a real-time HTTPS endpoint or a Batch Transform for offline scoring. Model Monitor watches production for drift, and Pipelines wires the steps into a repeatable, re-runnable CI/CD workflow that can retrain automatically.

In an interview

SageMaker manages the ML lifecycle on AWS so you don't run the infrastructure. Data lives in S3; you prototype in Studio, run managed training jobs on GPUs that provision and tear down automatically, and register versioned models with an approval gate. You serve two ways: a real-time endpoint for low-latency HTTPS inference, or Batch Transform to score a dataset offline. Model Monitor detects drift and quality decay, and Pipelines automates process-train-evaluate-register-deploy so retraining is repeatable.

Production defaults

Endpoints
serverless or async for spiky traffic; real-time only for steady traffic you can justify
Spot
for training. Checkpoint so interruptions cost minutes, not the run
Registry
version every model with its metrics attached. A model without its eval score is unshippable
Cost alarm
on idle endpoints. An unused real-time endpoint bills all month

What breaks

  • Large bill with no traffic — Idle real-time endpoints. Use serverless inference or shut them down.
  • Training job fails at the end — Output not written to the expected path. Checkpoint to S3 throughout, not only at completion.

Watch it explained

Introduction to Amazon SageMaker — Amazon Web Services, 4:47

Related