All concepts

Mixture of Experts (MoE)

A router sends each token to a few expert FFNs — huge capacity at a fraction of the compute.

Transformers & LLMs · Advanced · ~8 min

In plain English

Instead of one huge feed-forward layer that every token goes through, keep many smaller ones and let a router send each token to just two of them.

Why it's worth your time

It's how models get far more parameters without a proportional increase in cost per token — the architecture behind most frontier-scale models.

If you remember three things

  • Total parameters are large; active parameters per token are small
  • A learned router picks the top-k experts per token
  • Memory scales with total; compute scales with active

Overview

In a Mixture-of-Experts layer, the dense feed-forward is replaced by many expert FFNs plus a router. Per token, the router picks the top-1 or top-2 experts; only those run. This decouples parameter count from compute: the model can be enormous while each token only pays for a couple of experts.

How it works

  1. A token arrives In an MoE layer, a router decides which experts should process each token.
  2. Route to top-k experts The gate picks the top-1 or top-2 experts per token — only those run (here, experts 2 and 3).
  3. Selected experts compute Each expert is its own FFN; only the chosen few fire, so compute per token stays low.
  4. Combine outputs The chosen experts' outputs are combined, weighted by the router's gate scores.
  5. Sparse but huge MoE gives enormous parameter counts with the FLOPs of a much smaller dense model.

In an interview

MoE replaces the feed-forward layer with many experts and a router. For each token the router selects the top-k experts (usually 1 or 2), only those execute, and their outputs are combined weighted by the gate. So total parameters can be huge while the FLOPs per token stay near a small dense model — sparse activation is the whole trick.

Production defaults

Routing
top-2 of 8–64 experts is the common configuration
Load balancing
an auxiliary loss is required, or the router collapses onto a few experts
Serving reality
you must hold ALL experts in memory even though you use two. Memory-heavy, compute-light

What breaks

  • A few experts get all the traffic — Router collapse. The load-balancing loss coefficient is too low.
  • Cheaper per token but doesn't fit on your GPU — Exactly the trade. MoE saves compute, not memory — plan capacity on total parameters.

Watch it explained

What is Mixture of Experts? — IBM Technology, 7:57

Related