MoE total vs active parameters: what memory do you need?
The total MoE parameter count includes every expert and shared model component. Active parameters describe the subset used for a token’s computation, including shared components. A model with few active parameters can still require a large amount of memory because different tokens can select different experts.
For deployment engineers, use total weights to estimate storage and active computation to understand arithmetic. Neither count alone predicts serving latency.
What the router selects
A mixture-of-experts model replaces selected feed-forward networks with several expert networks. A router chooses a small number of experts for each token and combines their outputs. Shared experts, when present, run for every token. Other components, including attention, remain part of the active computation.
Experts are not necessarily smaller than the feed-forward network in a comparable dense model. Not every layer must use experts: DeepSeek-V3 keeps its first three layers dense.
| Model | Total parameters | Active parameters per token | Expert arrangement |
|---|---|---|---|
| Mixtral 8x7B | About 47B | About 13B | Eight routed experts; selects two per MoE layer |
| DeepSeek-V3 | 671B | 37B | 256 routed plus one shared expert; selects eight routed experts per MoE layer |
The counts come from the Mixtral and DeepSeek-V3 papers. They do not mean that only the selected weights need storage for the whole service.
Size storage and execution separately
At two bytes per weight, 671 billion weights require approximately 1.34 TB of raw storage before cache state, metadata, and runtime buffers. Quantization can reduce that weight payload. Sharding can distribute it across devices; offloading changes data-transfer requirements. None of these is implied by the 37B active count.
Routing also affects runtime. A batch can send many tokens to one expert and few to another. Expert parallelism transfers token representations between devices, adding communication. A larger batch can improve matrix utilization while changing which experts must run.
Check the exact checkpoint, precision, expert placement, supported kernels, network topology, and peak memory. Benchmark prompt and decode phases at realistic concurrency. Include uneven routing workloads if they occur in your traffic. Use measured token latency and goodput to judge cost rather than treating active parameters as an end-to-end price formula.
Engineering Guide: mixture of experts also covers routing balance during training.