Durable object storage in buckets — the cheap, near-infinite data lake behind ML on AWS
A bucket you put files in, that never runs out of space and effectively never loses anything. Cheap to store, and you pay to move data out.
It's where every dataset, model artifact and document corpus in an AWS-based AI system actually lives.
Amazon S3 is object storage: you put objects (bytes plus metadata) into buckets and address each by a full-path key, over a flat, effectively unlimited namespace. It is engineered for 99.999999999% (eleven nines) durability by replicating across Availability Zones, and its storage classes and lifecycle rules let you tier data from hot to cold to control cost. Because it is cheap, durable, and API-addressable, S3 is the default data lake feeding SageMaker, Athena, and analytics.
S3 stores objects in buckets, each addressed by a key — it's a key-value store, not a real filesystem. It gives eleven nines of durability by replicating across AZs, offers storage classes from Standard to Glacier with lifecycle rules to auto-tier by age, and supports versioning to undo bad writes or deletes. Objects are private by default and secured with IAM and bucket policies. In ML it's the data lake: training data, model artifacts, and logs all live in S3.
Introduction to Amazon Simple Storage Service (S3) - Cloud Storage on AWS — Amazon Web Services, 3:17