All concepts
Feature Store
A shared system for defining, computing, serving, and monitoring ML features.
MLOps & LLMOps · Intermediate · ~8 min
In plain English
One place where a feature is defined once and used by both training and serving, so the two can't quietly disagree.
Why it's worth your time
Training/serving skew is one of the hardest production ML bugs to find, and this is the structural fix.
If you remember three things
- Same definition, offline and online
- Point-in-time correctness prevents label leakage
- It's mainly a consistency guarantee, not a speed feature
Overview
A feature store is shared infrastructure for defining, computing, serving, and monitoring ML features once and reusing them everywhere. Its central value is a dual offline/online store fed by the same transformation logic, so the features a model trains on exactly match the features it sees at inference.
How it works
- Start: Feature Definitions Teams define reusable features with owners and freshness expectations.
- Feature Definitions -> Offline Store Historical feature values train models and backfills.
- Offline Store -> Online Store Low-latency feature lookup serves production predictions.
- Online Store -> Registry Schemas, lineage, and versions keep train/serve consistent.
- Registry -> Model Serving The same feature logic feeds training and inference.
In an interview
A feature store centralizes feature definitions and serves them from two backends: an offline store (a warehouse) for training and backfills, and a low-latency online store (Redis, DynamoDB) for real-time inference. The point is train/serve consistency plus reuse — the same versioned logic feeds both paths, so teams stop re-implementing features and avoid skew.
Production defaults
- Adopt when
- several models share features, or you've been bitten by skew. Not on day one
- Point-in-time joins
- non-negotiable. Joining current values to historical labels leaks the future
- Freshness
- define and monitor per feature. A stale feature is a silent accuracy loss
What breaks
- Great offline, poor online — The classic skew symptom. Log serving features and compare against the training set directly.
- Suspiciously perfect training accuracy — Time-travel leakage in the join. The feature knew something it couldn't have known at prediction time.