All concepts
Mixture of Experts (MoE)
A router sends each token to a few expert FFNs — huge capacity at a fraction of the compute.
Transformers & LLMs · Advanced · ~8 min
In plain English
Instead of one huge feed-forward layer that every token goes through, keep many smaller ones and let a router send each token to just two of them.
Why it's worth your time
It's how models get far more parameters without a proportional increase in cost per token — the architecture behind most frontier-scale models.
If you remember three things
- Total parameters are large; active parameters per token are small
- A learned router picks the top-k experts per token
- Memory scales with total; compute scales with active
Overview
In a Mixture-of-Experts layer, the dense feed-forward is replaced by many expert FFNs plus a router. Per token, the router picks the top-1 or top-2 experts; only those run. This decouples parameter count from compute: the model can be enormous while each token only pays for a couple of experts.
How it works
- A token arrives In an MoE layer, a router decides which experts should process each token.
- Route to top-k experts The gate picks the top-1 or top-2 experts per token — only those run (here, experts 2 and 3).
- Selected experts compute Each expert is its own FFN; only the chosen few fire, so compute per token stays low.
- Combine outputs The chosen experts' outputs are combined, weighted by the router's gate scores.
- Sparse but huge MoE gives enormous parameter counts with the FLOPs of a much smaller dense model.
In an interview
MoE replaces the feed-forward layer with many experts and a router. For each token the router selects the top-k experts (usually 1 or 2), only those execute, and their outputs are combined weighted by the gate. So total parameters can be huge while the FLOPs per token stay near a small dense model — sparse activation is the whole trick.
Production defaults
- Routing
- top-2 of 8–64 experts is the common configuration
- Load balancing
- an auxiliary loss is required, or the router collapses onto a few experts
- Serving reality
- you must hold ALL experts in memory even though you use two. Memory-heavy, compute-light
What breaks
- A few experts get all the traffic — Router collapse. The load-balancing loss coefficient is too low.
- Cheaper per token but doesn't fit on your GPU — Exactly the trade. MoE saves compute, not memory — plan capacity on total parameters.