Understanding MoE Models: Mixture of Experts for Local AI
What This Article Covers
- What MoE (Mixture of Experts) is technically, explained clearly and concisely.
- Why MoE matters for local AI: massive models, fast inference.
- The hardware trap: all parameters must fit in RAM/VRAM.
- Which MoE models make sense locally and what hardware they require.
Introduction
MoE (Mixture of Experts) powers some of the strongest open models in recent years: Mixtral, DeepSeek-V3/R1, Qwen3-MoE, Llama 4 Scout, gpt-oss-120b. Instead of routing every token through all parameters, an MoE model activates only a small subset of “experts” per token. A router network decides which ones.
The result: model capacity of a giant, with the speed of a compact model and that’s exactly why MoE matters for local AI.
How MoE Works (Briefly)
Dense model (classic): Every token passes through all layers with all parameters. A 70B model uses around 70B parameters per token, leading to slowness and computational intensity.
MoE model: The FFN layers are split into many “experts” (e.g., 128 of them). A router selects the top 2 to 8 matching experts per token.
Example Qwen3-30B-A3B: 30.5B parameters total, but only around 3.3B active per token. Inference speed like a 3B model, quality like something much larger.
Token → Router → selects 2-8 of 128 experts → only active ones computed
Why MoE Matters for Local AI
| Advantage | Explanation |
|---|---|
| Fast | Only active parameters are computed → more tokens/second |
| Quality | Large total parameter count = more knowledge stored |
| Efficient | Better quality/speed ratio than dense models |
| The Trap | Explanation |
|---|---|
| Total Memory | ALL parameters must fit in RAM/VRAM, including inactive ones |
| VRAM Demand | gpt-oss-120b needs around 60-80 GB, Qwen3-235B over 130 GB |
| Offloading | MoE plus RAM offloading (CPU/GPU mix) is trickier than with dense models |
Key MoE Models for Local Use
| Model | Total | Active | VRAM/RAM Required | Use Case |
|---|---|---|---|---|
| gpt-oss-120b | around 120B | around 5B active | around 60-80 GB | Best open model, agents, reasoning |
| Qwen3-30B-A3B | 30B | 3B | around 20 GB | All-rounder, fast, solid quality |
| Qwen3-235B-A22B | 235B | 22B | 130+ GB | Top quality, servers only |
| Mixtral 8x7B | 47B | around 13B | around 26 GB | Classic, well-established |
| Mixtral 8x22B | 141B | around 39B | 80+ GB | High quality, difficult locally |
| DeepSeek-V3/R1 | 671B | 37B | 400+ GB | Not realistic locally (API only) |
| Llama 4 Scout | 109B | 17B | around 60 GB | MoE with solid coding ability |
Hardware Reality
Here’s the crux: MoE models are fast, but memory-hungry. All parameters must be available, even if only a few are active per token.
- gpt-oss-120b (around 60-80 GB): Impossible on an RTX 4090 (24 GB). On a mini PC with 128 GB unified memory (MS-S1 Max) it runs, slow (around 10-15 tok/s), but completely local. See Homelab Guide.
- Qwen3-30B-A3B (around 20 GB): Runs on a 24-GB GPU or 32-GB Mac, the sweet spot MoE for desktop.
- CPU Inference: MoE on CPU is possible but slow. High memory bandwidth helps (see Memory Bandwidth). Unified-memory machines (Mac, Strix Halo) perform better than classical CPU+RAM setups.
MoE vs. Dense, When to Choose What
Choose MoE if:
- You have plenty of RAM/VRAM (>24 GB) and want maximum quality per token.
- Speed matters (active parameters are small).
- You want large models locally that would be prohibitively expensive as dense models.
Choose Dense if:
- Memory is tight (under 16 GB), a 7B/8B dense model is smaller and more predictable.
- Simplicity counts: dense models are straightforward to quantize and deploy.
- Latency is critical on weak hardware: dense has constant activation costs.
Practical Tips
- Ollama:
ollama run gpt-oss:120borqwen3:30b-a3b: MoE runs like any other model, just with higher RAM requirements. - Quantization: MoE models quantize well (Q4_M and so on): storage needs drop significantly.
- Offloading: With tight GPU memory, offload MoE layers to CPU (
--n-gpu-layers), see GPU Offloading. - llama.cpp: Supports MoE with expert parallelism, useful for multi-node clusters.
Further Reading
- IRC-Coding.de: Programming tutorials: model inference, GPU config.
- Homelab Guide: Hardware for large MoE models.
- AI Mini PC: 128-GB machines.
- Text Models: Model comparison.
- VRAM Calculator: Calculate memory requirements.
- Quantization: GGUF and quants.
Key Takeaways:
- MoE activates only some experts per token, fast like small, smart like large.
- But: all parameters must fit in RAM/VRAM, so total size determines memory needs.
- gpt-oss-120b, Qwen3-MoE, Mixtral are the relevant local MoE models.
- 128-GB machines (MS-S1 Max) make 120B MoE feasible locally.
- For under 16 GB hardware, dense models remain more practical.
FAQ
What is MoE?
Why is MoE faster?
What is the MoE trap?
Which MoE for home use?
MoE on CPU?
Is MoE the future?
Sources and Further Reading
- Mixtral Paper: MoE fundamentals.
- Qwen3: Qwen MoE models.
- IRC-Coding.de: Programming tutorials.


