Mixture of Experts

A model design that stores many sub-networks, called experts, and runs only a few of them per input, so capacity grows faster than compute cost.

13 min read

Β· Also in

By Ravi SuranaUpdated 13 sources

Quick answer

~20 sec

Mixture of Experts is a neural network design that splits one layer into many small sub-networks, called experts. A small gating network picks a few of them for each piece of input and skips the rest. So the model can hold far more parameters than it runs each time, and capacity grows faster than the cost of running it.

012 min

What an expert in MoE actually is

The word expert in Mixture of Experts is a technical name for a sub-network. It does not mean a person, and it does not mean a specialist in a subject. In a transformer language model, each layer contains a feed-forward block. That block transforms one position's representation. An MoE layer stores several copies of it, each copy with its own weights. Each copy is one expert. Nothing during training assigns a copy to biology or to code.

Readers assume otherwise so often that the Mixtral team measured it. Jiang and colleagues at Mistral AI ran Mixtral 8x7B over separate subsets of The Pile, a public text collection. They recorded which experts the model selected on each subset (arXiv:2401.04088, January 2024). The pattern of expert choices for arXiv papers looked much like the pattern for biology abstracts and for philosophy papers.

"Surprisingly, we do not observe obvious patterns in the assignment of experts based on the topic."

Jiang et al., *Mixtral of Experts*, arXiv:2401.04088, January 2024

What the same analysis did find was structural rather than topical. Indentation characters in Python code were sent to the same expert every time. Neighbouring words were often sent to the same expert too. On arXiv text at layer 15, the first-choice expert repeated for 27.9% of adjacent token pairs. Random choice across eight experts would give 12.5% (Table 5).

So experts do become specialised. They specialise in patterns below the level a person would give a name to.

022 min

The problem MoE solves

In a dense model, every parameter is used for every input. Dense here means no part of the network is skipped. Double the parameters and you roughly double the arithmetic for each unit of input, during training and then for every request the model ever serves. Capacity and running cost are locked to each other.

Shazeer and colleagues at Google Brain stated the limit directly in 2017. "The capacity of a neural network to absorb information is limited by its number of parameters" (arXiv:1701.06538, January 2017). More knowledge needs more parameters. In a dense model, more parameters needs more compute per input, permanently.

The idea for breaking that link already had a name: conditional computation, where only part of the network runs for any given input. The same paper describes its status at the time. Conditional computation "has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation. In practice, however, there are significant algorithmic and performance challenges."

Three challenges had blocked it. Accelerators are fast at large dense matrix operations and slow at small scattered ones, so skipping work does not automatically save time. The gating decision is discrete β€” an expert is either used or not β€” and discrete choices do not fit gradient-based training, which needs smooth quantities to adjust. And when experts are stored on different machines, sending each unit of input to its chosen machine costs network bandwidth, which can exceed the compute saved.

MoE is the design that answers those three at once.

032 min

How MoE routing works

Follow one short message through an MoE layer. Someone types reset my password into a chat box.

The route one token takes

First the text is split into tokens. A token is a chunk of text, usually a word or part of a word, and the model processes one token at a time. Say this message becomes four tokens. Follow the last one, password.

By the time password reaches an MoE layer it is no longer text. It is a vector: a list of a few thousand numbers. Those numbers hold what the model has worked out about that position so far, including the two words before it.

The router reads that vector. The router, also called the gating network, is one small learned matrix. It multiplies the vector and produces one score per expert. Mixtral 8x7B has eight experts in each MoE layer, so eight scores.

num_experts: 8 Β· top_k_experts: 2 β€” Mixtral 8x7B model architecture, Table 1, arXiv:2401.04088

The layer keeps the two highest scores and discards the other six. This is the sparse part. Sparse activation means most of the network performs no arithmetic for this token. The two surviving scores are converted into weights that add up to 1. The vector for password is then passed through those two experts only. Their outputs are multiplied by the two weights, added together, and the result is added back into the token's representation.

Then the same thing happens again at the next layer, decided again from scratch. The two experts chosen for password at layer 15 need not be the two chosen at layer 0. And reset, one position earlier, gets its own decision. Mixtral's authors state this plainly: the selected experts "can be different at each timestep".

Why the router can be trained

The weighting step is where the useful trick is. Each expert's output is multiplied by its gate weight before being added back, so the training signal reaches the router itself. The router is told, indirectly, whether the experts it picked helped.

That matters because picking the top two is an on/off choice, and gradient-based training cannot adjust an on/off choice. The weights on the winners, though, are continuous numbers, and adjusting those is enough for the router to learn. Shazeer and colleagues made this work at scale in 2017 by adding random noise to the scores before the selection (arXiv:1701.06538). The noise gives experts that would otherwise never win a chance to be picked and trained.

042 min

Total vs active parameters in MoE

Two parameter counts describe every MoE model, and confusing them is the most common practical error with this concept.

Total parameters is every weight stored. It sets how much memory the model needs. Active parameters is the subset that actually runs for one token. It sets how much arithmetic each token costs.

ModelPublishedTotal parametersActive per tokenExperts per MoE layerUsed per token
Switch Transformer (Switch-C)Jan 20211.6Tβ€”2,0481
Mixtral 8x7BJan 202447B13B82
OLMoE-1B-7BSep 20247B1B648
DeepSeek-V3Dec 2024671B37B256 routed + 1 shared8 routed + 1 shared

Rows in order: arXiv:2101.03961, arXiv:2401.04088, arXiv:2409.02060, arXiv:2412.19437. Switch-C's active count is blank because that paper does not quote it as one figure.

The name "8x7B" is worth correcting here, because it invites a wrong sum. Mixtral 8x7B is not eight 7B models, and 8 Γ— 7 does not give 47. Only the feed-forward blocks are duplicated into experts. Attention layers and embeddings are shared by all eight, so the total is lower than the multiplication suggests.

DeepSeek-V3 shows how far the two counts can separate. It stores 671B parameters and runs 37B for each token, which is about 5.5% of the model (DeepSeek-AI, arXiv:2412.19437, December 2024). Each of its MoE layers holds one shared expert, used by every token, and 256 routed experts, of which eight are selected.

052 min

Why MoE needs load balancing

Left alone, a router does not spread work evenly, and the imbalance reinforces itself. An expert chosen more often receives more training updates. It improves faster than the others. The router then scores it higher and sends it more tokens still. A few experts end up doing most of the work, and the others are stored, paid for in memory, and rarely used.

That costs in three ways:

  • Quality. Idle experts are capacity the model never gets to use.
  • Speed. Experts are usually stored on different accelerators, an arrangement called expert parallelism. Mixtral's authors note that this "introduces challenges in load balancing, as it is essential to distribute the workload evenly across the GPUs to prevent overloading individual GPUs" (arXiv:2401.04088). One overloaded device makes every other device wait.
  • Dropped tokens. Each expert gets a fixed buffer per batch, sized by a setting called the capacity factor. Tokens that arrive after the buffer is full skip that expert entirely (Fedus, Zoph and Shazeer, JMLR 2022). A smaller buffer is cheaper and drops more.

Three published answers exist, and they differ in where the correction is applied.

Shazeer's 2017 layer added an auxiliary loss. That is an extra training penalty which grows when the load is uneven, so the model is trained to balance as well as to predict. Switch Transformer kept that approach.

Zhou and colleagues at Google inverted the selection instead. In expert choice routing, each expert picks its own fixed number of tokens, rather than each token picking experts. Balance is then exact by construction. They reported more than 2x faster training convergence than Switch's top-1 gating and GShard's top-2 gating (arXiv:2202.09368, February 2022).

DeepSeek-V3 removed the penalty. An auxiliary loss strong enough to force balance also damages the prediction the model is being trained for. So DeepSeek-AI gave each expert a bias value, added to its routing score only. After every training step the bias of an overloaded expert is decreased and the bias of an underloaded one is increased. The weight multiplying the expert's output still comes from the original score (arXiv:2412.19437, December 2024).

062 min

What MoE costs to train and serve

Sparse activation reduces arithmetic. It does not reduce memory. Every expert has to be stored and loaded whether or not a given token uses it. Mixtral's paper states the consequence. Inference compute is set by the 13B active parameters, but "the memory costs for serving Mixtral are proportional to its sparse parameter count, 47B" (arXiv:2401.04088). Serving that model at the compute cost of a 13B one still needs hardware that can hold 47B parameters.

Three costs follow from that.

  • Memory. Sized by total parameters, never by active ones.
  • Utilisation. Routing adds work, and sending different tokens to different experts breaks each expert's job into smaller pieces. Mixtral's authors say MoE layers "are more suitable for batched workloads". One request at a time uses the hardware poorly.
  • Engineering. Expert parallelism, specialised kernels such as Megablocks, and continuous load monitoring are all machinery a dense model does not need.

The training side is where the saving is visible. DeepSeek-AI reported that DeepSeek-V3, at 671B total and 37B active, took 2.788M H800 GPU hours for its full training. The report prices that at US$5.576M, assuming a rental price of US$2 per GPU hour (Table 1, arXiv:2412.19437, December 2024). The same report says that figure covers only the official training run and excludes prior research and ablation experiments. So it is a floor for a model of that size, not the total cost of the programme that produced it.

072 min

What MoE means for your work

Four decisions change once you know which of the two parameter counts governs what.

An engineer choosing serving hardware sizes memory from the total count, not the active one. A 47B-parameter MoE will not fit where a 13B dense model fits, even though it costs roughly the same arithmetic per token. If the hard limit is the memory on one device, an MoE is the wrong shape. The right question is then a smaller dense model, or a quantised one.

A founder comparing inference cost per token should compare active parameters, not the number in the model's name. A provider serving a 671B-total, 37B-active model is doing roughly the per-token arithmetic of a 37B dense model. That is why such a model can be priced near a mid-size dense model while scoring against much larger ones. Comparing on the 671B figure makes the published price look impossible and leads to the wrong conclusion about what it costs to run.

An engineer sizing a workload should check batch size before assuming the saving appears. The efficiency depends on enough tokens arriving together for each expert to have work to do. An internal tool handling one request at a time will not see the advantage the active-parameter count implies.

Anyone planning a fine-tune should budget for more uncertainty than a dense model needs. Zoph and colleagues at Google named "training instabilities and uncertain quality during fine-tuning" as the specific obstacles that had kept sparse models from wider use (arXiv:2202.08906, February 2022). Their paper is a design guide for fixing exactly that. The fixes are known, and they are work somebody has to do.

082 min

MoE vs ensembles, routers and agents

One question separates MoE from all three of its neighbours: at what level is the choice made, and were the pieces trained together?

What runsHow it was trainedWhat decidesLevel of the decision
Mixture of Expertsa few sub-networks inside each layer of one modelexperts and router trained together, end to enda learned gating networkper token, per layer
Ensembleevery member model, then the outputs are combinedeach model trained separatelynothing decides; all are usednone β€” all run
Model routingone complete model per requestmodels trained independently, router trained separatelya separate router modelper request
Multi-agent systemseveral complete model calls in sequenceno joint training at allorchestration code or a promptper task or per step

Ensembles run every member and combine the answers, so cost rises with each member added. MoE runs a subset, so cost per token stays flat as experts are added. Those are opposite trades.

Model routing happens outside the models. RouteLLM, from Ong and colleagues, sends each query to either a stronger or a weaker complete model. They reported cost reductions "by over 2 times in certain cases" without a quality drop (arXiv:2406.18665, June 2024). The pieces there are separate products with separate endpoints. An MoE expert is a block of weights inside one file and cannot be called on its own.

Multi-agent systems are prompts and orchestration wrapped around complete models. Nothing in them is learned by gradient descent, and none of it changes a parameter count.

092 min

Where the MoE evidence is contested

The strongest published objection is not that MoE fails. It is that its advantage may shrink as models get larger.

Clark and colleagues at DeepMind fitted scaling laws across routed language models from 15M to 200B parameters, using up to 512 experts (arXiv:2202.01169, February 2022). Their first result supports the design: "Routing improves the performance of language models across all sizes and variants attempted." Their fitted law then adds a qualification. All three routing techniques they studied "predict diminishing improvements from routing when increasing scale". The gain per expert gets smaller as the base model grows.

They also put a number on where the gain would run out, and that number is extrapolated rather than measured. Routing, they wrote, "will remain beneficial up to models with base model size greater than 900 billion parameters". That is a prediction fitted from models up to 200B, not an observed limit, and as of September 2026 it is over four years old. It is the best available estimate, not a settled result.

Two further findings cut against comfortable explanations of the design.

Zoph and colleagues reported that training instability and unpredictable fine-tuning quality, rather than weak results, were what had kept sparse models out of wide use (arXiv:2202.08906). The obstacle was operational.

And Mistral AI's own measurement of which experts Mixtral selects found no topic structure at all, which contradicts the explanation most often offered for why the design works.

One caution is specific to this concept. Expert counts, parameter totals and routing details for several widely used commercial models are reported by third parties rather than published by the vendor. Those figures are unconfirmed. Do not plan pricing or capacity against them.

102 min

How MoE changed since 1991

The name comes from 1991, and the property that now defines the design has since reversed.

Jacobs, Jordan, Nowlan and Hinton published "Adaptive Mixtures of Local Experts" in Neural Computation 3(1), pages 79–87, in 1991. In their own description, the procedure "divides up a vowel discrimination task into appropriate subtasks, each of which can be solved by a very simple expert network". Every expert ran on every case, and the gating network blended their outputs. Nothing was skipped. Nothing was sparse.

Sparsity arrived in 2017. Shazeer and colleagues introduced the sparsely-gated MoE layer, placed between stacked LSTM layers. It held up to thousands of experts and as many as 137B parameters. They reported "greater than 1000x improvements in model capacity with only minor losses in computational efficiency" (arXiv:1701.06538).

Then it moved into transformers. Lepikhin and colleagues trained a 600B-parameter sparsely-gated MoE transformer for translation from 100 languages into English, on 2,048 TPU v3 chips in 4 days (GShard, arXiv:2006.16668, June 2020). Fedus, Zoph and Shazeer then cut routing to a single expert per token. They reported up to 7x faster pre-training than T5-Base and T5-Large, and scaled to 1.6 trillion parameters with 2,048 experts (arXiv:2101.03961, January 2021; JMLR 2022).

The 2024 change was granularity. Dai and colleagues split experts into smaller pieces, and reserved some as shared experts that every token uses. At 145B scale they reported results comparable to DeepSeek 67B at 28.5% of the computation (DeepSeekMoE, arXiv:2401.06066, January 2024).

So "mixture" is now the least accurate word in the name. In 1991 the outputs of all experts were mixed together. Today most of them never run.

112 min

What MoE does not solve

MoE buys one thing: more parameters at a fixed compute cost per token. Everything outside that is unchanged.

  • It does not reduce memory. If a model does not fit on the device you have, adding experts makes the fit worse. Quantisation, distillation, or a smaller dense model address that problem. This design does not.
  • It does not add knowledge. Capacity is room for knowledge, not knowledge itself. DeepSeek-V3's 671B parameters were pre-trained on 14.8 trillion tokens (arXiv:2412.19437). Without that data, the extra experts would have nothing to hold.
  • It does not make a model explainable. There is no single expert you can open to see what the model knows about medicine.
  • It does not remove the latency floor on one request. At a batch size of one, the weights of the selected experts still have to be read out of memory. The routing step adds work of its own.

The overclaim worth correcting is the equivalence claim. An MoE with a very large total parameter count is not the same thing as a dense model of that size. The honest comparison is the kind the papers actually publish. Mixtral, at 47B total and 13B active, "outperforms or matches Llama 2 70B and GPT-3.5 across all evaluated benchmarks" (arXiv:2401.04088, January 2024). That is a claim against named models, on named benchmarks, at a stated date. A claim of the form "671B parameters, therefore as capable as any 671B model" has nothing behind it.

?8 questions

Questions people ask

Is Mixture of Experts the same as an ensemble?

No. An ensemble runs every model it contains and combines the answers, so cost grows with each model added. An MoE runs only a few of its experts per token, so adding experts raises memory but not per-token compute. An ensemble's members are also trained separately; an MoE's experts and router are trained together as one model.

Do MoE experts specialise in subjects like law or code?

Generally no. Mistral AI measured expert selection in Mixtral 8x7B across topic-separated subsets of The Pile and reported no obvious topic pattern (arXiv:2401.04088, January 2024). Expert choice tracked syntax and position instead, such as indentation characters in code going to the same expert.

Does MoE reduce the memory needed to run a model?

No. It reduces arithmetic per token, not storage. All experts must be held in memory because any of them may be selected. Mixtral's paper states that serving memory is proportional to its 47B total parameters, while inference compute follows its 13B active parameters (arXiv:2401.04088).

What are active parameters in an MoE model?

Active parameters are the weights that actually run for a single token, as opposed to the total weights stored. DeepSeek-V3 stores 671B parameters and activates 37B per token (arXiv:2412.19437, December 2024). Active parameters predict compute cost and price per token; total parameters predict memory.

Why does Mixtral 8x7B have 47B parameters instead of 56B?

Because only the feed-forward blocks are duplicated into eight experts. Attention layers and embeddings are shared across all experts, so the name's multiplication overstates the total. The published figures are 47B total and 13B active per token, with two of eight experts used (arXiv:2401.04088).

What is load balancing in MoE and why does it matter?

Load balancing is any mechanism that stops the router from sending most tokens to a few experts. Without it the imbalance grows on its own, because a frequently chosen expert trains faster and is then chosen more. Imbalance wastes capacity, idles accelerators, and causes tokens to be dropped when an expert's per-batch buffer fills.

Does MoE still help at very large scale?

It is contested. Clark and colleagues at DeepMind found routing improved performance at every size they tried, up to 200B parameters. But their fitted scaling law predicts shrinking gains as models grow, and estimates benefit persisting only up to roughly 900B base parameters (arXiv:2202.01169, February 2022). That figure is extrapolated, not measured.

Is MoE the same as routing between different LLMs?

No. Routing between LLMs, as in RouteLLM (arXiv:2406.18665, June 2024), picks one complete, independently trained model per request. MoE picks sub-networks inside a single model, fresh at every layer and for every token, and its experts cannot be called separately.

Β§13 sources

Sources on MoE

  1. Jacobs, Jordan, Nowlan & Hinton, "Adaptive Mixtures of Local Experts", 1991

  2. Shazeer et al., "Outrageously Large Neural Networks", 2017

  3. Lepikhin et al., "GShard", 2020

  4. Fedus, Zoph & Shazeer, "Switch Transformers", 2021

Show all 13 sources
  1. Switch Transformers, JMLR 23, 2022

  2. Clark et al. (DeepMind), "Unified Scaling Laws for Routed Language Models", 2022

  3. Zoph et al., "ST-MoE", 2022

  4. Zhou et al., "Mixture-of-Experts with Expert Choice Routing", 2022

  5. Jiang et al. (Mistral AI), "Mixtral of Experts", 2024

  6. Dai et al., "DeepSeekMoE", 2024

  7. DeepSeek-AI, "DeepSeek-V3 Technical Report", 2024

  8. Ong et al., "RouteLLM", 2024

  9. Muennighoff et al., "OLMoE", 2024

Keep reading

More from AI

All of AI
All of AI