Buying RAM for Local AI
What this article covers
- Why RAM matters for local AI.
- How much RAM different model sizes require.
- Differences between DDR4, DDR5, and ECC.
- Why dual-channel improves performance.
- Buying tips and common mistakes.
Introduction: Buying RAM for Local AI
RAM is one of the two most critical resources for local AI models, right alongside VRAM. When models don’t fit in GPU memory, they spill into system RAM, either partially or entirely. Larger, faster RAM means smoother inference. Choosing the right RAM setup for Ollama, llama.cpp, or agent deployments eliminates bottlenecks and load times.
This article helps you pick the right amount and type of RAM for your local AI projects.
Key terms
- RAM: System memory.
- VRAM: GPU memory.
- DDR4: Common, affordable RAM standard.
- DDR5: Newer standard with higher bandwidth.
- ECC: Error-Correcting Code, catches and fixes bit errors.
- Dual-Channel: Two RAM sticks in parallel for more bandwidth.
- Memory bandwidth: Speed at which data flows through RAM.
- Unified Memory: RAM and VRAM shared as one pool, like on Apple Silicon.
Why does RAM matter?
Local AI models contain billions of parameters that must stay in memory for the model to generate responses. During CPU inference or when a model exceeds GPU memory, the entire model sits in RAM. The model also maintains a KV-cache in memory during processing, which grows with context length.
Rule of thumb for model sizes
Using 4-bit quantization as a rough baseline:
| Model size | RAM needed |
|---|---|
| 7B / 8B | 6 to 8 GB |
| 13B / 14B | 10 to 12 GB |
| 30B | 20 to 24 GB |
| 70B | 40 to 48 GB |
These numbers cover only the model itself. Your operating system, Ollama, databases, and other services consume additional RAM. Plan for some headroom.
Recommended minimum amounts
- Starting out with 7B models: 16 GB RAM.
- Comfortable use with 13B/14B: 32 GB RAM.
- 30B models: 64 GB RAM.
- 70B models: 96 to 128 GB RAM.
Extra RAM almost always pays off. It lets you run larger models, support longer contexts, or host multiple services at once.
DDR4 vs. DDR5
DDR4
- Cheaper and widely available.
- Good for existing systems and budget builds.
- Sufficient for most local AI setups.
DDR5
- Higher bandwidth.
- Better performance during CPU inference.
- More expensive, especially at larger capacities.
- Worth it if CPU inference is your focus.
For pure GPU inference, the gap between DDR4 and DDR5 narrows considerably, since the model lives in VRAM. Memory bandwidth becomes more important for CPU inference.
ECC RAM
- Error correction for long-running calculations.
- Valuable for servers and homelabs operating 24/7.
- Available on many server platforms and workstation CPUs.
- Not mandatory for consumer test machines.
If you run a 24/7 server, ECC offers real benefits. Most desktop PCs with consumer-grade AMD or Intel CPUs don’t support ECC.
Dual-Channel
Two matching RAM sticks running in dual-channel mode deliver significantly more memory bandwidth than a single stick. You’ll notice the difference during CPU inference. When upgrading RAM, always buy pairs and install them in the correct slots.
Apple Silicon and Unified Memory
On Apple Silicon Macs, the CPU and GPU share a single memory pool. 16 GB works for small 7B models, while 32 to 64 GB of unified memory handles larger ones. The bandwidth is very high, making Apple Silicon solid for AI work even without a dedicated GPU.
Memory bandwidth
Bandwidth affects CPU inference performance. DDR4-3200 delivers roughly 51 GB/s per channel, while DDR5-4800 provides about 76 GB/s per channel. Dual-channel doubles these figures. Models with many parameters benefit significantly from higher bandwidth.
Buying recommendations
Budget build for 7B models
- 32 GB DDR4-3200, dual-channel.
- Enough for Ollama with 8B models.
Mid-range for 14B models
- 64 GB DDR4 or DDR5.
- Supports larger models and RAG services.
Workstation for 30B to 70B models
- 128 GB DDR4/DDR5, optionally ECC.
- Ideal for CPU inference or as a buffer when VRAM runs short.
Common buying mistakes
- Not enough RAM: Models get swapped to disk or crash.
- Single stick: You miss dual-channel performance gains.
- Slow RAM timings: Reduces effective bandwidth.
- Incompatible ECC: Motherboard or CPU doesn’t support it.
- Mismatched modules: Different sticks can cause stability issues.
- Underestimating Apple Silicon: Unified memory works differently than separate RAM and VRAM.
Further resources
- BotServ.de GPU for local AI
- BotServ.de Ollama Performance
- BotServ.de Sizing hardware correctly
- BotServ.de RAM and VRAM requirements
FAQ: RAM for local AI
Do I need more RAM than VRAM? Yes, if your model exceeds GPU memory, it lives in RAM. The operating system and other services also consume RAM.
Is 16 GB enough for Ollama? For small 7B models, yes. For larger models or RAG setups, no.
Is DDR5 significantly better for AI? For CPU inference, yes. For pure GPU inference, less relevant.
Is ECC worth it for a homelab? Yes, if your server runs continuously and stability matters.
Should I buy 32 GB DDR4 or 16 GB DDR5? Generally, 32 GB DDR4 is better since more RAM enables larger models.
Sources and further reading
- Ollama GitHub: https://github.com/ollama/ollama
- DDR5 overview: https://www.crucial.de/articles/about-memory/memory-types
Summary: Buying RAM for Local AI
RAM is central to local AI, especially for CPU inference or large models. 16 to 32 GB suffices for 7B models, while 64 GB or more makes sense for larger ones. DDR5 offers more bandwidth, DDR4 costs less. Dual-channel matters. ECC helps with stable, long-running servers. When you match RAM capacity and type to your model size and use case, Ollama and other local AI tools deliver noticeably smoother results.


