Mixture of Experts · Language Models · Efficient AI

Mixture of Experts Explained: Sparse Routing from Switch to DeepSeek-V3

MoE replaces one big feed-forward layer with many experts plus a router, so a model stores far more parameters than it spends per token. Switch Transformer, Mixtral 8x7B and DeepSeek-V3 mark the arc.

Mixture of Experts Explained: Sparse Routing from Switch to DeepSeek-V3

How it works

A standard Transformer runs every token through the same feed-forward network at every layer. A Mixture-of-Experts (MoE) model replaces that single block with N parallel copies, the “experts,” plus a small router network that decides, per token and per layer, which experts handle it. Only the chosen experts run, so the model stores far more parameters than it spends compute on. That is the whole deal: parameter count stops being a proxy for per-token cost.

Three papers mark the arc of the idea in modern LLMs, and each solved a different bottleneck:

  • Switch Transformer (Google, 2021) made sparse MoE trainable at scale by betting that routing each token to one expert is enough. It scaled to 1.6 trillion parameters while holding per-token FLOPs near a T5-Base dense model, and hit up to 7x faster pretraining than T5-Base/Large at the same compute.
  • Mixtral 8x7B (Mistral, 2023) turned MoE into something the open community could download and serve: 8 experts, top-2 routing, 47B total parameters but ~13B active per token, matching or beating Llama 2 70B and GPT-3.5 under Apache 2.0.
  • DeepSeek-V3 (2024) pushed sparsity to its current frontier: 671B parameters, 37B active per token, trained on 14.8 trillion tokens for 2.788M H800 GPU hours (about $5.6M at the report’s $2/GPU-hour assumption), with a routing setup that balances experts without a quality-eroding auxiliary loss.

Key numbers

DimensionSwitch TransformerMixtral 8x7BDeepSeek-V3
Total parameters1.6T (Switch-C)47B671B
Active parameters per tokenFLOPs held near T5-Base~13B37B
Active share of parameterslow (single expert fires)~28% (13B of 47B)~5.5% (37B of 671B)
Routing choice per tokentop-1, exactly one experttop-2 of 8 expertsa small subset of experts per token
Load balancingauxiliary loss + capacity factor, overflowing tokens droppednot the paper’s focusauxiliary-loss-free, via a per-expert routing bias
Headline training resultup to 7x faster pretraining than T5-Base/Large at equal computematches or beats Llama 2 70B and GPT-3.5 with 13B-class computefrontier-class scores for roughly an order of magnitude less compute than assumed
Licenseresearch releaseApache 2.0open weights

Read the active-share row twice, because it is the part people consistently get wrong. Mixtral runs about 28% of its parameters per token; DeepSeek-V3 runs about 5.5%. Both are “MoE models,” but they sit at very different points on the sparse-dense spectrum, and their engineering problems differ accordingly.

Why Switch went to one expert, and why the field walked back

Before Switch, the assumption inherited from early MoE work was top-k routing with k ≥ 2: send each token to at least two experts so the router gets a useful gradient. Switch’s bet was that k = 1 is enough: route to the single highest-scoring expert and scale its output by the router probability so gradients still flow. The payoff was immediate. Routing compute halves, cross-device communication roughly halves, and the layer becomes far simpler to implement, because every token has exactly one destination.

The result held up: up to 7x faster pretraining than T5-Base and T5-Large at identical compute, and Switch-C reached 1.6T parameters. But two of Switch’s mechanisms did not age as well. The capacity factor, a fixed buffer per expert with overflow tokens simply dropped through the residual, means some tokens get no expert at all. And the field’s later verdict on k = 1 is visible in what shipped afterwards: Mixtral returned to top-2, and DeepSeek-V3 routes to a small subset rather than a single expert. Single-expert routing was simple and fast, but not the quality-optimal point. Switch’s lasting contributions are the framing (sparsity as a way to grow capacity without growing per-token cost), the load-balancing problem it made everyone solve, and selective float32 router precision in an otherwise bfloat16 model, which made large sparse training stable in low precision for the first time.

What Mixtral actually is, and what “8x7B” gets wrong

Mixtral keeps the Mistral 7B architecture but replaces each layer’s feed-forward block with 8 of them. A router scores every token’s hidden state over the 8 experts, runs the top-2, and combines their outputs with the router’s softmax weights. The selection is per token, per layer. The same sentence can send consecutive tokens to entirely different expert pairs, and a token’s experts at layer 5 say nothing about its experts at layer 20.

The name “8x7B” misleads in both directions. The experts share the attention layers, so the total is 47B, not 56B. And only the feed-forward path is sparse, so about 13B parameters are active per token, not the full 47B. The marketing pitch, 47B quality at 13B inference cost, is real, with the honest catch that it is a server-side bargain: routing is dynamic, so you cannot know which experts a request will need in advance, and all 47B must be resident in memory. Mixtral is cheap in compute and throughput, not in VRAM. It needs the memory of a 47B model to run at the speed of a 13B one, which is why single-consumer-GPU users almost immediately reach for quantization.

Mixtral’s evaluation story is Mistral’s own: with 13B-class active compute it matched or beat Llama 2 70B and GPT-3.5 on every benchmark Mistral ran. That made it the first strong open MoE the community could actually serve, and the Apache 2.0 license did as much work as the architecture: MoE serving stacks, fine-tunes, and quantizations proliferated within weeks. One caveat from the paper’s own analysis: the router does not assign experts to human-interpretable topics. “Experts” is an architectural label, not a semantic one; you cannot steer the model by picking an expert for math or code.

How DeepSeek-V3 changed the routing problem

DeepSeek-V3 is best read as a cost-engineering execution of an already-validated design. Of its 671B parameters, a router fires only 37B per token, so you pay a 671B model’s capacity at roughly a 37B model’s per-token FLOPs. Sparsity is extreme, about 5.5% of parameters active, which means the load-balancing problem Switch solved with an auxiliary loss becomes acute: if the router favors a few experts, the other 95% of your capacity sits idle.

The auxiliary-loss-free fix is the paper’s signature idea, and it is worth understanding precisely because it inverts the usual trade. Auxiliary balancing losses work but fight the main training objective, quietly degrading quality. DeepSeek-V3 instead nudges balance by adjusting a per-expert bias term inside the routing computation, with no competing loss term at all. Experts stay busy without taxing the thing you care about. Around that, V3 wraps two more efficiency primitives: Multi-head Latent Attention, which shrinks the inference KV cache, and multi-token prediction training, which densifies supervision and enables speculative decoding.

The headline is the bill: 14.8 trillion training tokens in 2.788 million H800 GPU hours, roughly $5.6M at the report’s assumed $2 per GPU-hour, for a model that matches leading closed-source systems on many benchmarks. Before V3, “frontier-class needs a closed lab’s budget” was the unstated premise of the field. V3 made that premise look negotiable, and because DeepSeek-R1 was built on V3, its efficiency directly underwrote the reasoning model that rattled the closed labs.

MoE vs dense: what you actually trade

The comparison that matters is not parameters but constraints:

  • Compute per token. MoE wins clearly. Mixtral bills like a ~13B model; DeepSeek-V3 like a ~37B model, while their capacity is that of far larger networks. If you serve high traffic and FLOPs are your bill, sparsity is the best deal in LLM architecture.
  • Memory. MoE loses clearly, and this is the part the “runs like a smaller model” pitch hides. All experts must be resident because routing is input-dependent: 47B of weights for Mixtral, 671B for DeepSeek-V3. Dense models waste compute but never memory.
  • Serving economics. MoE shines under batching, where expert loads even out across many concurrent requests and per-token cost approaches the active-parameter count. Under a single-stream, latency-bound workload, the memory cost is paid but the throughput dividend is not collected.
  • Quality per dollar of training. Switch showed sparse models reach a given quality up to 7x faster at fixed compute; DeepSeek-V3 showed the same logic at frontier scale for roughly an order of magnitude less assumed cost.
  • Operational complexity. Expert parallelism, load balancing, token dropping in Switch’s design, and routing instability in low precision are all real engineering burdens that dense models simply do not have.

Failure modes worth knowing

Three recur across the three papers. First, router collapse: without balancing pressure, a few experts hoard tokens and the rest idle, and the entire capacity argument evaporates. Switch answered with an auxiliary loss and capacity buffers; DeepSeek-V3 with a bias term; both are answers to the same disease. Second, the memory illusion: because only a fraction of parameters fire per token, people assume MoE shrinks deployment footprint. It does not. Mixtral needs 47B of VRAM to deliver its 13B speed, and the gap is widest exactly where MoE’s compute win is largest. Third, interpretability theater: neither Mixtral’s nor Switch’s experts map to human concepts; routing is learned for performance, so you cannot debug or steer a model through its experts. Anyone selling “pick the math expert” is describing an architecture label, not an observed behavior.

Limits and open questions

Everything in the table above comes from each lab’s own report, and the evaluations are not cross-comparable: Switch’s speedups are wall-clock-to-quality against T5, Mixtral’s wins are Mistral-run benchmarks against 2023 models, and DeepSeek-V3’s cost figure assumes a $2/GPU-hour rate that is a modeling choice, not a market price. No paper here runs a dense and a sparse model of matched total quality through the same harness at matched training cost, so “MoE is more efficient” rests on each lab’s internal comparisons rather than a controlled experiment. And the active-parameter counts measure feed-forward routing only; attention, which Switch shares and DeepSeek-V3 additionally optimizes with MLA, follows different economics entirely.

FAQ

What is a mixture of experts model?

A Transformer variant where each layer’s feed-forward network is replaced by many parallel copies (“experts”) plus a router. For every token, the router picks a small subset of experts to run, so the model stores many more parameters than it spends compute on per token. The three canonical examples are Switch Transformer (one expert per token, 1.6T parameters), Mixtral 8x7B (2 of 8 experts, 47B total, ~13B active), and DeepSeek-V3 (37B of 671B active per token).

Is MoE better than a dense model?

For capacity per unit of compute, yes; that is the entire point, and it held from Switch’s 7x pretraining speedup to DeepSeek-V3’s frontier-class model at roughly $5.6M of reported training cost. For memory, no: every expert must stay loaded, so a 47B MoE needs 47B of VRAM even though each token touches ~13B. MoE is a compute bargain and a memory liability; dense is the reverse.

Why does Mixtral have 47B parameters and not 56B?

Because the 8 experts only replace the feed-forward blocks; attention layers are shared across all experts. Counting shared parameters once instead of eight times gives 47B. Likewise, only ~13B parameters are active per token because only 2 of the 8 experts run for each token, and only in the feed-forward path.

Why did the field abandon Switch Transformer’s one-expert routing?

Switch showed k = 1 trains well and is fast, but later systems returned to routing each token to at least two experts (Mixtral’s top-2 of 8) or a small subset (DeepSeek-V3). Single-expert routing was the simplest point in the design space, not the quality-optimal one. What survived from Switch is the framing (sparsity as capacity without per-token cost) and the load-balancing machinery everyone still builds on.

How does DeepSeek-V3 balance experts without an auxiliary loss?

It adjusts a per-expert bias term in the routing computation to keep loads even, instead of adding a separate balancing loss that competes with the main training objective. The distinction matters because auxiliary losses measurably trade model quality for balance; the bias approach claims to get balance without that tax.

Do MoE models save VRAM?

No. Routing depends on the input, so you cannot know which experts a request will need; all experts must be resident. Mixtral runs at ~13B-class compute speed but needs 47B of weights in memory, and DeepSeek-V3’s 37B active compute comes with a 671B memory footprint. MoE trades memory for FLOPs, in the direction opposite to what the “runs like a small model” slogan suggests.