012 min
What an expert in MoE actually is
The word expert in Mixture of Experts is a technical name for a sub-network. It does not mean a person, and it does not mean a specialist in a subject. In a transformer language model, each layer contains a feed-forward block. That block transforms one position's representation. An MoE layer stores several copies of it, each copy with its own weights. Each copy is one expert. Nothing during training assigns a copy to biology or to code.
Readers assume otherwise so often that the Mixtral team measured it. Jiang and colleagues at Mistral AI ran Mixtral 8x7B over separate subsets of The Pile, a public text collection. They recorded which experts the model selected on each subset (arXiv:2401.04088, January 2024). The pattern of expert choices for arXiv papers looked much like the pattern for biology abstracts and for philosophy papers.
"Surprisingly, we do not observe obvious patterns in the assignment of experts based on the topic."
What the same analysis did find was structural rather than topical. Indentation characters in Python code were sent to the same expert every time. Neighbouring words were often sent to the same expert too. On arXiv text at layer 15, the first-choice expert repeated for 27.9% of adjacent token pairs. Random choice across eight experts would give 12.5% (Table 5).
So experts do become specialised. They specialise in patterns below the level a person would give a name to.
022 min
The problem MoE solves
In a dense model, every parameter is used for every input. Dense here means no part of the network is skipped. Double the parameters and you roughly double the arithmetic for each unit of input, during training and then for every request the model ever serves. Capacity and running cost are locked to each other.
Shazeer and colleagues at Google Brain stated the limit directly in 2017. "The capacity of a neural network to absorb information is limited by its number of parameters" (arXiv:1701.06538, January 2017). More knowledge needs more parameters. In a dense model, more parameters needs more compute per input, permanently.
The idea for breaking that link already had a name: conditional computation, where only part of the network runs for any given input. The same paper describes its status at the time. Conditional computation "has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation. In practice, however, there are significant algorithmic and performance challenges."
Three challenges had blocked it. Accelerators are fast at large dense matrix operations and slow at small scattered ones, so skipping work does not automatically save time. The gating decision is discrete β an expert is either used or not β and discrete choices do not fit gradient-based training, which needs smooth quantities to adjust. And when experts are stored on different machines, sending each unit of input to its chosen machine costs network bandwidth, which can exceed the compute saved.
MoE is the design that answers those three at once.
032 min
How MoE routing works
Follow one short message through an MoE layer. Someone types reset my password into a chat box.
The route one token takes
First the text is split into tokens. A token is a chunk of text, usually a word or part of a word, and the model processes one token at a time. Say this message becomes four tokens. Follow the last one, password.
By the time password reaches an MoE layer it is no longer text. It is a vector: a list of a few thousand numbers. Those numbers hold what the model has worked out about that position so far, including the two words before it.
The router reads that vector. The router, also called the gating network, is one small learned matrix. It multiplies the vector and produces one score per expert. Mixtral 8x7B has eight experts in each MoE layer, so eight scores.
num_experts: 8Β·top_k_experts: 2β Mixtral 8x7B model architecture, Table 1, arXiv:2401.04088
The layer keeps the two highest scores and discards the other six. This is the sparse part. Sparse activation means most of the network performs no arithmetic for this token. The two surviving scores are converted into weights that add up to 1. The vector for password is then passed through those two experts only. Their outputs are multiplied by the two weights, added together, and the result is added back into the token's representation.
Then the same thing happens again at the next layer, decided again from scratch. The two experts chosen for password at layer 15 need not be the two chosen at layer 0. And reset, one position earlier, gets its own decision. Mixtral's authors state this plainly: the selected experts "can be different at each timestep".
Why the router can be trained
The weighting step is where the useful trick is. Each expert's output is multiplied by its gate weight before being added back, so the training signal reaches the router itself. The router is told, indirectly, whether the experts it picked helped.
That matters because picking the top two is an on/off choice, and gradient-based training cannot adjust an on/off choice. The weights on the winners, though, are continuous numbers, and adjusting those is enough for the router to learn. Shazeer and colleagues made this work at scale in 2017 by adding random noise to the scores before the selection (arXiv:1701.06538). The noise gives experts that would otherwise never win a chance to be picked and trained.
042 min
Total vs active parameters in MoE
Two parameter counts describe every MoE model, and confusing them is the most common practical error with this concept.
Total parameters is every weight stored. It sets how much memory the model needs. Active parameters is the subset that actually runs for one token. It sets how much arithmetic each token costs.
| Model | Published | Total parameters | Active per token | Experts per MoE layer | Used per token |
|---|---|---|---|---|---|
| Switch Transformer (Switch-C) | Jan 2021 | 1.6T | β | 2,048 | 1 |
| Mixtral 8x7B | Jan 2024 | 47B | 13B | 8 | 2 |
| OLMoE-1B-7B | Sep 2024 | 7B | 1B | 64 | 8 |
| DeepSeek-V3 | Dec 2024 | 671B | 37B | 256 routed + 1 shared | 8 routed + 1 shared |
Rows in order: arXiv:2101.03961, arXiv:2401.04088, arXiv:2409.02060, arXiv:2412.19437. Switch-C's active count is blank because that paper does not quote it as one figure.
The name "8x7B" is worth correcting here, because it invites a wrong sum. Mixtral 8x7B is not eight 7B models, and 8 Γ 7 does not give 47. Only the feed-forward blocks are duplicated into experts. Attention layers and embeddings are shared by all eight, so the total is lower than the multiplication suggests.
DeepSeek-V3 shows how far the two counts can separate. It stores 671B parameters and runs 37B for each token, which is about 5.5% of the model (DeepSeek-AI, arXiv:2412.19437, December 2024). Each of its MoE layers holds one shared expert, used by every token, and 256 routed experts, of which eight are selected.
052 min
Why MoE needs load balancing
Left alone, a router does not spread work evenly, and the imbalance reinforces itself. An expert chosen more often receives more training updates. It improves faster than the others. The router then scores it higher and sends it more tokens still. A few experts end up doing most of the work, and the others are stored, paid for in memory, and rarely used.
That costs in three ways:
- Quality. Idle experts are capacity the model never gets to use.
- Speed. Experts are usually stored on different accelerators, an arrangement called expert parallelism. Mixtral's authors note that this "introduces challenges in load balancing, as it is essential to distribute the workload evenly across the GPUs to prevent overloading individual GPUs" (arXiv:2401.04088). One overloaded device makes every other device wait.
- Dropped tokens. Each expert gets a fixed buffer per batch, sized by a setting called the capacity factor. Tokens that arrive after the buffer is full skip that expert entirely (Fedus, Zoph and Shazeer, JMLR 2022). A smaller buffer is cheaper and drops more.
Three published answers exist, and they differ in where the correction is applied.
Shazeer's 2017 layer added an auxiliary loss. That is an extra training penalty which grows when the load is uneven, so the model is trained to balance as well as to predict. Switch Transformer kept that approach.
Zhou and colleagues at Google inverted the selection instead. In expert choice routing, each expert picks its own fixed number of tokens, rather than each token picking experts. Balance is then exact by construction. They reported more than 2x faster training convergence than Switch's top-1 gating and GShard's top-2 gating (arXiv:2202.09368, February 2022).
DeepSeek-V3 removed the penalty. An auxiliary loss strong enough to force balance also damages the prediction the model is being trained for. So DeepSeek-AI gave each expert a bias value, added to its routing score only. After every training step the bias of an overloaded expert is decreased and the bias of an underloaded one is increased. The weight multiplying the expert's output still comes from the original score (arXiv:2412.19437, December 2024).
062 min
What MoE costs to train and serve
Sparse activation reduces arithmetic. It does not reduce memory. Every expert has to be stored and loaded whether or not a given token uses it. Mixtral's paper states the consequence. Inference compute is set by the 13B active parameters, but "the memory costs for serving Mixtral are proportional to its sparse parameter count, 47B" (arXiv:2401.04088). Serving that model at the compute cost of a 13B one still needs hardware that can hold 47B parameters.
Three costs follow from that.
- Memory. Sized by total parameters, never by active ones.
- Utilisation. Routing adds work, and sending different tokens to different experts breaks each expert's job into smaller pieces. Mixtral's authors say MoE layers "are more suitable for batched workloads". One request at a time uses the hardware poorly.
- Engineering. Expert parallelism, specialised kernels such as Megablocks, and continuous load monitoring are all machinery a dense model does not need.
The training side is where the saving is visible. DeepSeek-AI reported that DeepSeek-V3, at 671B total and 37B active, took 2.788M H800 GPU hours for its full training. The report prices that at US$5.576M, assuming a rental price of US$2 per GPU hour (Table 1, arXiv:2412.19437, December 2024). The same report says that figure covers only the official training run and excludes prior research and ablation experiments. So it is a floor for a model of that size, not the total cost of the programme that produced it.
072 min
What MoE means for your work
Four decisions change once you know which of the two parameter counts governs what.
An engineer choosing serving hardware sizes memory from the total count, not the active one. A 47B-parameter MoE will not fit where a 13B dense model fits, even though it costs roughly the same arithmetic per token. If the hard limit is the memory on one device, an MoE is the wrong shape. The right question is then a smaller dense model, or a quantised one.
A founder comparing inference cost per token should compare active parameters, not the number in the model's name. A provider serving a 671B-total, 37B-active model is doing roughly the per-token arithmetic of a 37B dense model. That is why such a model can be priced near a mid-size dense model while scoring against much larger ones. Comparing on the 671B figure makes the published price look impossible and leads to the wrong conclusion about what it costs to run.
An engineer sizing a workload should check batch size before assuming the saving appears. The efficiency depends on enough tokens arriving together for each expert to have work to do. An internal tool handling one request at a time will not see the advantage the active-parameter count implies.
Anyone planning a fine-tune should budget for more uncertainty than a dense model needs. Zoph and colleagues at Google named "training instabilities and uncertain quality during fine-tuning" as the specific obstacles that had kept sparse models from wider use (arXiv:2202.08906, February 2022). Their paper is a design guide for fixing exactly that. The fixes are known, and they are work somebody has to do.
082 min
MoE vs ensembles, routers and agents
One question separates MoE from all three of its neighbours: at what level is the choice made, and were the pieces trained together?
| What runs | How it was trained | What decides | Level of the decision | |
|---|---|---|---|---|
| Mixture of Experts | a few sub-networks inside each layer of one model | experts and router trained together, end to end | a learned gating network | per token, per layer |
| Ensemble | every member model, then the outputs are combined | each model trained separately | nothing decides; all are used | none β all run |
| Model routing | one complete model per request | models trained independently, router trained separately | a separate router model | per request |
| Multi-agent system | several complete model calls in sequence | no joint training at all | orchestration code or a prompt | per task or per step |
Ensembles run every member and combine the answers, so cost rises with each member added. MoE runs a subset, so cost per token stays flat as experts are added. Those are opposite trades.
Model routing happens outside the models. RouteLLM, from Ong and colleagues, sends each query to either a stronger or a weaker complete model. They reported cost reductions "by over 2 times in certain cases" without a quality drop (arXiv:2406.18665, June 2024). The pieces there are separate products with separate endpoints. An MoE expert is a block of weights inside one file and cannot be called on its own.
Multi-agent systems are prompts and orchestration wrapped around complete models. Nothing in them is learned by gradient descent, and none of it changes a parameter count.
092 min
Where the MoE evidence is contested
The strongest published objection is not that MoE fails. It is that its advantage may shrink as models get larger.
Clark and colleagues at DeepMind fitted scaling laws across routed language models from 15M to 200B parameters, using up to 512 experts (arXiv:2202.01169, February 2022). Their first result supports the design: "Routing improves the performance of language models across all sizes and variants attempted." Their fitted law then adds a qualification. All three routing techniques they studied "predict diminishing improvements from routing when increasing scale". The gain per expert gets smaller as the base model grows.
They also put a number on where the gain would run out, and that number is extrapolated rather than measured. Routing, they wrote, "will remain beneficial up to models with base model size greater than 900 billion parameters". That is a prediction fitted from models up to 200B, not an observed limit, and as of September 2026 it is over four years old. It is the best available estimate, not a settled result.
Two further findings cut against comfortable explanations of the design.
Zoph and colleagues reported that training instability and unpredictable fine-tuning quality, rather than weak results, were what had kept sparse models out of wide use (arXiv:2202.08906). The obstacle was operational.
And Mistral AI's own measurement of which experts Mixtral selects found no topic structure at all, which contradicts the explanation most often offered for why the design works.
One caution is specific to this concept. Expert counts, parameter totals and routing details for several widely used commercial models are reported by third parties rather than published by the vendor. Those figures are unconfirmed. Do not plan pricing or capacity against them.
102 min
How MoE changed since 1991
The name comes from 1991, and the property that now defines the design has since reversed.
Jacobs, Jordan, Nowlan and Hinton published "Adaptive Mixtures of Local Experts" in Neural Computation 3(1), pages 79β87, in 1991. In their own description, the procedure "divides up a vowel discrimination task into appropriate subtasks, each of which can be solved by a very simple expert network". Every expert ran on every case, and the gating network blended their outputs. Nothing was skipped. Nothing was sparse.
Sparsity arrived in 2017. Shazeer and colleagues introduced the sparsely-gated MoE layer, placed between stacked LSTM layers. It held up to thousands of experts and as many as 137B parameters. They reported "greater than 1000x improvements in model capacity with only minor losses in computational efficiency" (arXiv:1701.06538).
Then it moved into transformers. Lepikhin and colleagues trained a 600B-parameter sparsely-gated MoE transformer for translation from 100 languages into English, on 2,048 TPU v3 chips in 4 days (GShard, arXiv:2006.16668, June 2020). Fedus, Zoph and Shazeer then cut routing to a single expert per token. They reported up to 7x faster pre-training than T5-Base and T5-Large, and scaled to 1.6 trillion parameters with 2,048 experts (arXiv:2101.03961, January 2021; JMLR 2022).
The 2024 change was granularity. Dai and colleagues split experts into smaller pieces, and reserved some as shared experts that every token uses. At 145B scale they reported results comparable to DeepSeek 67B at 28.5% of the computation (DeepSeekMoE, arXiv:2401.06066, January 2024).
So "mixture" is now the least accurate word in the name. In 1991 the outputs of all experts were mixed together. Today most of them never run.
112 min
What MoE does not solve
MoE buys one thing: more parameters at a fixed compute cost per token. Everything outside that is unchanged.
- It does not reduce memory. If a model does not fit on the device you have, adding experts makes the fit worse. Quantisation, distillation, or a smaller dense model address that problem. This design does not.
- It does not add knowledge. Capacity is room for knowledge, not knowledge itself. DeepSeek-V3's 671B parameters were pre-trained on 14.8 trillion tokens (arXiv:2412.19437). Without that data, the extra experts would have nothing to hold.
- It does not make a model explainable. There is no single expert you can open to see what the model knows about medicine.
- It does not remove the latency floor on one request. At a batch size of one, the weights of the selected experts still have to be read out of memory. The routing step adds work of its own.
The overclaim worth correcting is the equivalence claim. An MoE with a very large total parameter count is not the same thing as a dense model of that size. The honest comparison is the kind the papers actually publish. Mixtral, at 47B total and 13B active, "outperforms or matches Llama 2 70B and GPT-3.5 across all evaluated benchmarks" (arXiv:2401.04088, January 2024). That is a claim against named models, on named benchmarks, at a stated date. A claim of the form "671B parameters, therefore as capable as any 671B model" has nothing behind it.
?8 questions
Questions people ask
Is Mixture of Experts the same as an ensemble?
Do MoE experts specialise in subjects like law or code?
Does MoE reduce the memory needed to run a model?
What are active parameters in an MoE model?
Why does Mixtral 8x7B have 47B parameters instead of 56B?
What is load balancing in MoE and why does it matter?
Does MoE still help at very large scale?
Is MoE the same as routing between different LLMs?
Β§13 sources
Sources on MoE
Jacobs, Jordan, Nowlan & Hinton, "Adaptive Mixtures of Local Experts", 1991
Shazeer et al., "Outrageously Large Neural Networks", 2017
Lepikhin et al., "GShard", 2020
Fedus, Zoph & Shazeer, "Switch Transformers", 2021
Show all 13 sourcesShow fewer sources
Switch Transformers, JMLR 23, 2022
Clark et al. (DeepMind), "Unified Scaling Laws for Routed Language Models", 2022
Zoph et al., "ST-MoE", 2022
Zhou et al., "Mixture-of-Experts with Expert Choice Routing", 2022
Jiang et al. (Mistral AI), "Mixtral of Experts", 2024
Dai et al., "DeepSeekMoE", 2024
DeepSeek-AI, "DeepSeek-V3 Technical Report", 2024
Ong et al., "RouteLLM", 2024
Muennighoff et al., "OLMoE", 2024





