Choosing a model
MoE models: total vs active parameters
By llamaperf · · 6 min read
Quick answer
Size memory by the total parameter count and expect generation speed closer to a model the size of the active count. That makes MoE models a good fit for machines with lots of memory and modest bandwidth, and it's why keeping the experts in system RAM works better than you'd expect.
Two parameter counts
Mixture-of-experts models come with two sizes in their names, like 30B-A3B: 30 billion parameters in total, 3 billion active for each token. Both numbers matter, for different reasons.
Total parameters set the memory
Each MoE layer holds many expert blocks, and a small router picks a few of them for every token. Any expert might be picked next, so all of them have to be loaded. The Hugging Face MoE explainer is blunt about it:
However, all parameters need to be loaded in RAM, so memory requirements are high. From Mixture of Experts Explained on Hugging Face
Active parameters set the work per token
Only the chosen experts do any work for a given token, plus the parts every token shares, such as attention. Mistral described its Mixtral model this way:
Concretely, Mixtral has 46.7B total parameters but only uses 12.9B parameters per token. From Mistral's Mixtral announcement
So you pay for a 47B model in memory and roughly a 13B model in work per token. The Hugging Face post explains why it's 47B and not eight times 7B: only the feed-forward layers are split into experts, and the rest is shared.
Working out the memory
Treat an MoE model like a dense model of its total size when you plan memory. The active count doesn't help you here.
Use the total count
Here's the arithmetic for a 46.7B model at Q4_K_M, which averages about 4.9 bits per weight (see our post on GGUF quant names): 46.7 billion times 4.9 bits, divided by eight, comes to about 28.6 GB. That's the download and the memory for the weights alone, whatever the active count says.
Add the context
The KV cache depends on the attention layers, and those are shared, so an MoE model's cache is sized like any other model with the same attention shape. Our post on context length and the KV cache walks through the maths. Put both parts into the VRAM calculator, which reads the total and active counts separately for sparse models, and check the fit with the context you really use.
Why speed tracks the active count
Generating a token on a single-user setup is mostly limited by how fast your hardware can read the weights it needs. An MoE model only reads the active experts, so it reads far fewer bytes per token than a dense model of the same total size.
Fewer bytes read per token
Take the Mixtral figures above as an example. At about 4.9 bits per weight, the 12.9B active parameters come to roughly 7.9 GB read per token, against about 28.6 GB if all 46.7B were read. On the same memory bandwidth, that's more than three times fewer bytes per token. That's why a big MoE model can feel quick on a machine with lots of fairly slow memory, like a Mac with unified memory or a PC with a large amount of system RAM.
Where the estimate breaks
The picture changes when many tokens are processed together. Reading a long prompt, or serving several users at once, sends different tokens to different experts, so far more of the model gets read in each step. Prompt processing on an MoE model is therefore closer to the cost of its total size than its active size. Router overhead and uneven expert use also shave a bit off the ideal. Estimates are a starting point, and the GPU reports show what people measure in practice.
Keeping the experts in system RAM
Because each token touches only a few experts, you can keep the expert weights in system RAM and the shared layers on the GPU. The GPU handles the parts every token needs, and the CPU reads just the few experts each token picks.
llama.cpp's --cpu-moe option
The llama.cpp server documentation lists --cpu-moe to "keep all Mixture of Experts (MoE) weights in the CPU", and --n-cpu-moe N to do that for only the first N layers. With the second one you can fill your VRAM with as many experts as fit and leave the rest in RAM. That's more targeted than plain layer offload, which moves whole layers and their shared parts along with them. Our post on partial offload covers the layer-based approach.
Test it on your machine
How well this works depends on your system RAM speed and how many experts each token uses. Start with --cpu-moe, measure generation speed and time to first token on a real prompt, then lower the expert count on the CPU with --n-cpu-moe until VRAM is nearly full. Compare the result against a dense model that fits entirely on your GPU. If you're on a Mac, the Mac reports show how MoE models run on unified memory, where there's no split to tune.
Frequently asked questions
Do I need to fit the whole MoE model in VRAM?
All the experts must be in memory, but that memory can be split. You can keep the shared layers on the GPU and the expert weights in system RAM with llama.cpp's --cpu-moe or --n-cpu-moe options. It runs slower than all-GPU and much faster than you might expect.
How much VRAM does an MoE model take compared to a dense model?
The same as a dense model with the same total parameter count at the same quantization. A 46.7B MoE at Q4_K_M needs about 28.6 GB for the weights, like any 46.7B model. The active count only affects speed.
Why is MoE generation faster than a dense model of the same size?
Each token only uses the experts the router picks, so the hardware reads far fewer weights per token. Mixtral uses 12.9B of its 46.7B parameters per token, so it reads roughly a quarter of the bytes a dense 46.7B model would.
How fast does an MoE run if only the active parameters fit in VRAM?
It depends on your system RAM bandwidth, because the experts are read from RAM on every token. Measure it on your own machine with the experts on the CPU and compare against a dense model that fits entirely on the GPU.
Is a 30B MoE smarter than a 32B dense model?
Not by default. Quality depends on the particular model. The MoE is usually faster. Run both on the tasks you care about and compare the answers along with the speed.
Does prompt processing get the same speedup on an MoE model?
Less of it. A long prompt sends different tokens to different experts, so most of the model gets read while the prompt is processed. Expect prompt processing to cost closer to the total parameter count.