AWS's managed ML platform: train on GPUs, version models, and serve them as endpoints
Managed infrastructure for the whole model lifecycle — training jobs, tuning, and endpoints — so you're not building GPU orchestration yourself.
It's the default answer to 'how do we train and serve this on AWS' and shows up in every AWS-flavoured system design round.
SageMaker is Amazon's managed platform for the whole ML lifecycle. You build in Studio notebooks, launch training jobs that spin up a managed GPU cluster on demand and tear it down when done, version artifacts in the Model Registry, and deploy either as a real-time HTTPS endpoint or a Batch Transform for offline scoring. Model Monitor watches production for drift, and Pipelines wires the steps into a repeatable, re-runnable CI/CD workflow that can retrain automatically.
SageMaker manages the ML lifecycle on AWS so you don't run the infrastructure. Data lives in S3; you prototype in Studio, run managed training jobs on GPUs that provision and tear down automatically, and register versioned models with an approval gate. You serve two ways: a real-time endpoint for low-latency HTTPS inference, or Batch Transform to score a dataset offline. Model Monitor detects drift and quality decay, and Pipelines automates process-train-evaluate-register-deploy so retraining is repeatable.
Introduction to Amazon SageMaker — Amazon Web Services, 4:47