All concepts

AWS S3

Durable object storage in buckets — the cheap, near-infinite data lake behind ML on AWS

Cloud Services · Beginner · ~4 min

In plain English

A bucket you put files in, that never runs out of space and effectively never loses anything. Cheap to store, and you pay to move data out.

Why it's worth your time

It's where every dataset, model artifact and document corpus in an AWS-based AI system actually lives.

If you remember three things

  • Object storage, not a filesystem — no real directories
  • Storage classes trade retrieval time and cost
  • Egress is the line item that surprises people

Overview

Amazon S3 is object storage: you put objects (bytes plus metadata) into buckets and address each by a full-path key, over a flat, effectively unlimited namespace. It is engineered for 99.999999999% (eleven nines) durability by replicating across Availability Zones, and its storage classes and lifecycle rules let you tier data from hot to cold to control cost. Because it is cheap, durable, and API-addressable, S3 is the default data lake feeding SageMaker, Athena, and analytics.

In an interview

S3 stores objects in buckets, each addressed by a key — it's a key-value store, not a real filesystem. It gives eleven nines of durability by replicating across AZs, offers storage classes from Standard to Glacier with lifecycle rules to auto-tier by age, and supports versioning to undo bad writes or deletes. Objects are private by default and secured with IAM and bucket policies. In ML it's the data lake: training data, model artifacts, and logs all live in S3.

Production defaults

Layout
prefix by date or entity for parallel reads. A million objects in one flat prefix is a performance problem
Lifecycle
rules to move cold data to cheaper classes automatically
Formats
Parquet over CSV for analytics — smaller, typed, column-pruned
Security
block public access at the account level, encrypt at rest, least-privilege bucket policies

What breaks

  • Bill dominated by data transfer — Egress, not storage. Keep compute in the same region as the data.
  • Listing a bucket is slow — Too many objects under one prefix. Partition the key space.

Watch it explained

Introduction to Amazon Simple Storage Service (S3) - Cloud Storage on AWS — Amazon Web Services, 3:17

Related