Explainer
MoE Models and VRAM: Why Total Parameters Decide Memory, Not Active Ones
Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset
A model card that says '109B total, 17B active' raises an obvious question: do I need memory for 109 billion parameters or for 17? The answer is 109. This page explains why, and what the smaller number is good for.
Total and active parameters are handled for every model. Free.
Check a MoE model on your GPU →What a mixture-of-experts model is
In a standard, or dense, language model, every token passes through every weight in every layer.
In a mixture-of-experts (MoE) model, the feed-forward part of each layer is split into many separate blocks called experts. A small router looks at each token and sends it to a few of them, often two to eight out of dozens or hundreds. The other experts are skipped for that token.
Two numbers follow:
- Total parameters: everything in the model, all experts included.
- Active parameters: what is used to process one token: the shared parts plus the few chosen experts.
Which number decides memory
Total. The router chooses different experts for different tokens, and different ones again at each layer. Across one sentence, most experts get used. They all have to be in memory, ready.
So a model with 109B total parameters needs about as much memory as a dense 109B model, whatever its active count.
Which number decides speed
Active. For each token, the hardware reads and computes only the active parameters. A MoE model with 17B active parameters generates tokens at roughly the pace of a dense 17B model, provided the whole model fits in fast memory.
That is the appeal: the knowledge capacity of a very large model, at the generation speed of a much smaller one.
Examples from the id8 dataset
| Model | Parameters | VRAM at Q4_K_M, 8K context |
|---|---|---|
| Llama 4 Scout | 109B total, 17B active | 64.9 GB |
| Qwen 3 30B-A3B | 30B total, 3B active | 18.8 GB |
| Gemma 4 26B (MoE) | 25.2B total, 4B active | 15.8 GB |
| Qwen 3 32B (dense) | 32B | 21.2 GB |
Compare the two models of about 30B total parameters: one is MoE with a small active count and one is dense. Their memory needs are close, because total size is close. Their speed is not: the MoE model computes about a tenth as many parameters per token.
What this means when choosing a local model
If memory is your limit, MoE does not help. Look at total parameters.
If speed is your limit and you have the memory, MoE helps a great deal. A 30B MoE model with 3B active can be pleasant to use on hardware where a dense 30B model feels slow.
On a Mac with a lot of unified memory, MoE models are a particularly good fit. Macs have plenty of memory and less raw compute than a high-end GPU, and MoE needs exactly that combination.
With partial offload to system RAM, MoE models degrade more gracefully than dense ones, because less data is read per token. They are still much slower than when fully in VRAM.
Common misreadings
- "17B active means it runs on a 16 GB card." No. Memory follows the total.
- "MoE models are lower quality." Not as a rule. Many current open-weight leaders are MoE. Compare scores, not architecture.
- "Active parameters are the real size." They describe the compute per token, not the knowledge stored.
- "Quantization works differently." It works the same way. A Q4 MoE model is about a quarter of its full-precision size.
The context cache is separate
On top of the weights, the model keeps a cache for the tokens in the context. Its size depends on the attention layers and the context length, not on the number of experts. A long context adds the same kind of memory to a MoE model as to a dense one. See What is the KV cache?.
Fine-tuning MoE models
Fine-tuning needs all weights in memory plus training overhead, so again the total count applies. Check a specific model in Can I Fine-Tune LLM?.
Check a model
Can I Run LLM? shows total and active parameters for every MoE model and calculates memory from the total. Enter your GPU to see what fits.
Frequently asked questions
Do I need VRAM for total or active parameters?
Total. All experts must be loaded, because which ones are used changes with every token.
Then what do active parameters affect?
Speed. Only the active parameters are computed for each token, so a MoE model generates faster than a dense model of the same total size.
How much VRAM does Llama 4 Scout need?
Llama 4 Scout has 109B total, 17B active parameters. At Q4_K_M with an 8K context it needs about 64.9 GB.
Is a MoE model better than a dense model?
It is a different trade. For the same memory, a dense model computes all its weights for every token; a MoE model computes a fraction, which is faster. Quality depends on the specific models, so compare benchmark scores.
Can I load only the experts I need?
Not in practice. The router picks experts per token and per layer, so all of them are needed within a single sentence.