All concepts

Shadow Deployment

Run a new model on production traffic without serving its answers yet.

MLOps & LLMOps · Advanced · ~8 min

In plain English

Run the new model alongside the old one on real traffic, but throw its answers away. You get production evidence with zero production risk.

Why it's worth your time

It's how you find out that the model which won offline loses online — before any user sees it.

If you remember three things

  • Real traffic, real latency, no user impact
  • Compare distributions and disagreements, not just accuracy
  • Costs double inference for the shadow period

Overview

Shadow deployment mirrors live production traffic to a candidate model whose outputs are computed but never shown to users. It de-risks a launch by measuring the new model on real inputs and load, then compares latency, errors, cost, and disagreement against the incumbent before any user is exposed.

How it works

  1. Start: Production Traffic Real requests flow to the current model.
  2. Production Traffic -> Current Model The live model serves user-visible responses.
  3. Current Model -> Shadow Model A candidate model receives copied traffic but its output is hidden.
  4. Shadow Model -> Compare Measure latency, errors, safety, cost, and disagreement.
  5. Compare -> Promote or Block Only promote if shadow metrics beat release gates.

In an interview

In shadow mode, production requests are duplicated to a candidate model in parallel with the live one, but the candidate's responses are discarded, not served. You measure it on real traffic — latency, error rate, cost, safety, and disagreement with the incumbent — with zero user risk. It validates behavior under real load before you promote via canary or A/B.

Production defaults

Duration
long enough to cover a full traffic cycle — at least a week, including a weekend
Compare
prediction distributions, disagreement rate, and p95 latency under real load
Then
canary to a small traffic slice before full rollout. Shadow proves it runs; canary proves it helps

What breaks

  • Shadow looked fine, live rollout was bad — Shadow doesn't capture feedback loops — the new model's outputs change user behaviour. Canary does.
  • Shadow doubled your latency — It's running in the request path. Run it async, off the critical path.

Watch it explained

Top 5 Most-Used Deployment Strategies — ByteByteGo, 10:00

Related