All concepts

GCP Cloud Storage

Object storage buckets on GCP — the tiered, durable data lake for Vertex and BigQuery

Cloud Services · Beginner · ~4 min

In plain English

Google's version of the same bucket: put files in, they stay, you pay for space and for moving data out.

Why it's worth your time

Same role as S3 in a GCP-based system — the landing zone for every dataset and document corpus.

If you remember three things

  • Object storage with storage classes by access frequency
  • Integrates directly with BigQuery and Vertex AI
  • Egress and cross-region transfer are the cost traps

Overview

Google Cloud Storage (GCS) is object storage: buckets holding key-addressed objects in a flat namespace, GCP's equivalent of S3. A bucket's location (region, dual-region, or multi-region) fixes where copies live, storage classes from Standard to Archive plus Object Lifecycle rules tier data from hot to cold, and object versioning lets you undo mistakes. Access is governed by IAM with uniform bucket-level access, and GCS is the data lake that feeds Vertex AI training and BigQuery external tables.

In an interview

GCS stores objects in buckets, each named by a full-path key over a flat namespace — the folders are just naming. You choose a location (region, dual-, or multi-region) for locality versus redundancy, and storage classes (Standard, Nearline, Coldline, Archive) with Object Lifecycle rules to auto-tier and delete by age. Versioning recovers overwrites and deletes, IAM with uniform bucket-level access keeps objects private by default, and GCS is the lake feeding Vertex AI and BigQuery.

Production defaults

Location
keep buckets in the same region as the compute that reads them
Lifecycle
auto-transition cold objects to Nearline/Coldline
Formats
Parquet/Avro for analytics; BigQuery can query them in place
Security
uniform bucket-level access, CMEK where policy requires it

What breaks

  • Unexpected transfer costs — Cross-region reads. Co-locate storage and compute.
  • Slow reads from a training job — Many small files. Consolidate into larger shards sized for parallel reads.

Watch it explained

How to store data on Google Cloud — Google Cloud Tech, 6:52

Related