All concepts

Feature Store

A shared system for defining, computing, serving, and monitoring ML features.

MLOps & LLMOps · Intermediate · ~8 min

In plain English

One place where a feature is defined once and used by both training and serving, so the two can't quietly disagree.

Why it's worth your time

Training/serving skew is one of the hardest production ML bugs to find, and this is the structural fix.

If you remember three things

  • Same definition, offline and online
  • Point-in-time correctness prevents label leakage
  • It's mainly a consistency guarantee, not a speed feature

Overview

A feature store is shared infrastructure for defining, computing, serving, and monitoring ML features once and reusing them everywhere. Its central value is a dual offline/online store fed by the same transformation logic, so the features a model trains on exactly match the features it sees at inference.

How it works

  1. Start: Feature Definitions Teams define reusable features with owners and freshness expectations.
  2. Feature Definitions -> Offline Store Historical feature values train models and backfills.
  3. Offline Store -> Online Store Low-latency feature lookup serves production predictions.
  4. Online Store -> Registry Schemas, lineage, and versions keep train/serve consistent.
  5. Registry -> Model Serving The same feature logic feeds training and inference.

In an interview

A feature store centralizes feature definitions and serves them from two backends: an offline store (a warehouse) for training and backfills, and a low-latency online store (Redis, DynamoDB) for real-time inference. The point is train/serve consistency plus reuse — the same versioned logic feeds both paths, so teams stop re-implementing features and avoid skew.

Production defaults

Adopt when
several models share features, or you've been bitten by skew. Not on day one
Point-in-time joins
non-negotiable. Joining current values to historical labels leaks the future
Freshness
define and monitor per feature. A stale feature is a silent accuracy loss

What breaks

  • Great offline, poor online — The classic skew symptom. Log serving features and compare against the training set directly.
  • Suspiciously perfect training accuracy — Time-travel leakage in the join. The feature knew something it couldn't have known at prediction time.

Watch it explained

Introduction to Vertex AI Feature Store — Google Cloud Tech, 7:01

Related