Skip to content
BotServBotServ
MoEMixture of Expertsgpt-oss-120bMixtralQwen3DeepSeekModels

Understanding MoE Models: Mixture of Experts for Local AI

Learn what Mixture-of-Experts models are, why they matter for local AI, and which MoE models deliver results.

S

schutzgeist

4 min read
Understanding MoE Models: Mixture of Experts for Local AI

Understanding MoE Models: Mixture of Experts for Local AI

What This Article Covers

  • What MoE (Mixture of Experts) is technically, explained clearly and concisely.
  • Why MoE matters for local AI: massive models, fast inference.
  • The hardware trap: all parameters must fit in RAM/VRAM.
  • Which MoE models make sense locally and what hardware they require.

Introduction

MoE (Mixture of Experts) powers some of the strongest open models in recent years: Mixtral, DeepSeek-V3/R1, Qwen3-MoE, Llama 4 Scout, gpt-oss-120b. Instead of routing every token through all parameters, an MoE model activates only a small subset of “experts” per token. A router network decides which ones.

The result: model capacity of a giant, with the speed of a compact model and that’s exactly why MoE matters for local AI.

How MoE Works (Briefly)

Dense model (classic): Every token passes through all layers with all parameters. A 70B model uses around 70B parameters per token, leading to slowness and computational intensity.

MoE model: The FFN layers are split into many “experts” (e.g., 128 of them). A router selects the top 2 to 8 matching experts per token.

Example Qwen3-30B-A3B: 30.5B parameters total, but only around 3.3B active per token. Inference speed like a 3B model, quality like something much larger.

Token → Router → selects 2-8 of 128 experts → only active ones computed

Why MoE Matters for Local AI

AdvantageExplanation
FastOnly active parameters are computed → more tokens/second
QualityLarge total parameter count = more knowledge stored
EfficientBetter quality/speed ratio than dense models
The TrapExplanation
Total MemoryALL parameters must fit in RAM/VRAM, including inactive ones
VRAM Demandgpt-oss-120b needs around 60-80 GB, Qwen3-235B over 130 GB
OffloadingMoE plus RAM offloading (CPU/GPU mix) is trickier than with dense models

Key MoE Models for Local Use

ModelTotalActiveVRAM/RAM RequiredUse Case
gpt-oss-120baround 120Baround 5B activearound 60-80 GBBest open model, agents, reasoning
Qwen3-30B-A3B30B3Baround 20 GBAll-rounder, fast, solid quality
Qwen3-235B-A22B235B22B130+ GBTop quality, servers only
Mixtral 8x7B47Baround 13Baround 26 GBClassic, well-established
Mixtral 8x22B141Baround 39B80+ GBHigh quality, difficult locally
DeepSeek-V3/R1671B37B400+ GBNot realistic locally (API only)
Llama 4 Scout109B17Baround 60 GBMoE with solid coding ability

Hardware Reality

Here’s the crux: MoE models are fast, but memory-hungry. All parameters must be available, even if only a few are active per token.

  • gpt-oss-120b (around 60-80 GB): Impossible on an RTX 4090 (24 GB). On a mini PC with 128 GB unified memory (MS-S1 Max) it runs, slow (around 10-15 tok/s), but completely local. See Homelab Guide.
  • Qwen3-30B-A3B (around 20 GB): Runs on a 24-GB GPU or 32-GB Mac, the sweet spot MoE for desktop.
  • CPU Inference: MoE on CPU is possible but slow. High memory bandwidth helps (see Memory Bandwidth). Unified-memory machines (Mac, Strix Halo) perform better than classical CPU+RAM setups.

MoE vs. Dense, When to Choose What

Choose MoE if:

  • You have plenty of RAM/VRAM (>24 GB) and want maximum quality per token.
  • Speed matters (active parameters are small).
  • You want large models locally that would be prohibitively expensive as dense models.

Choose Dense if:

  • Memory is tight (under 16 GB), a 7B/8B dense model is smaller and more predictable.
  • Simplicity counts: dense models are straightforward to quantize and deploy.
  • Latency is critical on weak hardware: dense has constant activation costs.

Practical Tips

  • Ollama: ollama run gpt-oss:120b or qwen3:30b-a3b: MoE runs like any other model, just with higher RAM requirements.
  • Quantization: MoE models quantize well (Q4_M and so on): storage needs drop significantly.
  • Offloading: With tight GPU memory, offload MoE layers to CPU (--n-gpu-layers), see GPU Offloading.
  • llama.cpp: Supports MoE with expert parallelism, useful for multi-node clusters.

Further Reading

Key Takeaways:

  • MoE activates only some experts per token, fast like small, smart like large.
  • But: all parameters must fit in RAM/VRAM, so total size determines memory needs.
  • gpt-oss-120b, Qwen3-MoE, Mixtral are the relevant local MoE models.
  • 128-GB machines (MS-S1 Max) make 120B MoE feasible locally.
  • For under 16 GB hardware, dense models remain more practical.

FAQ

What is MoE?

Mixture of Experts: instead of using all parameters per token, a router network selects 2-8 expert subnets from many options. The result is large overall capacity with small compute costs per token.

Why is MoE faster?

Because only active experts are computed. Qwen3-30B-A3B uses only around 3B parameters per token: inference speed like a 3B model, but with 30B total knowledge.

What is the MoE trap?

All parameters must fit in memory, including inactive ones. gpt-oss-120b needs around 60-80 GB, even though only around 5B are active per token. MoE is fast but memory-hungry.

Which MoE for home use?

Qwen3-30B-A3B (around 20 GB) for typical GPU PCs; gpt-oss-120b for 128-GB machines like the MS-S1 Max. Mixtral 8x7B as a compact classic.

MoE on CPU?

Possible but slow: MoE benefits from memory bandwidth. On unified-memory machines (Mac, Strix Halo) it works better than on classic CPU+RAM. For CPU, a small dense model is usually faster.

Is MoE the future?

It looks that way: almost all new flagship models (DeepSeek, Qwen3, Llama 4, gpt-oss) are MoE. The architecture scales better than dense.

Sources and Further Reading

Back to Blog
Share:

Related Posts