Buying RAM for Local AI
What this article covers
- Why RAM matters for local AI.
- How much RAM different models need.
- DDR4 vs. DDR5 and clock speeds.
- ECC and dual-channel configurations.
- Practical recommendations.
Introduction: Buying RAM for local AI
Alongside the GPU and CPU, RAM is one of the most critical components for running local AI. If you want to work with larger models or handle long contexts, you’ll need substantial memory. Insufficient RAM forces the model to spill data onto slow disks or swap space, which tanks performance. The right RAM choice ensures models run fast and reliably.
This article explains how much RAM you need and what to look for when buying.
Key terms
- RAM: Random Access Memory.
- DDR4/DDR5: Memory standards.
- Clock speed: Speed measured in MHz.
- Timings: Latency values.
- ECC: Error-correcting code memory.
- Channels: Parallel memory banks.
- Dual-Channel: Two channels in parallel.
- Quad-Channel: Four channels in parallel.
- UDIMM, RDIMM: Memory modules for desktops and servers.
- Unbuffered: No buffering, lower latency.
How much RAM do I need?
| Model size | Quantization | RAM requirement |
|---|---|---|
| 3B | Q4 | 4-6 GB |
| 7B | Q4 | 8-12 GB |
| 13B | Q4 | 14-20 GB |
| 30B | Q4 | 30-40 GB |
| 70B | Q4 | 80-100 GB |
These figures assume CPU inference. For GPU inference, the model must fit in VRAM, while the remaining system RAM handles the OS, context, and other applications.
DDR4 vs. DDR5
| Aspect | DDR4 | DDR5 |
|---|---|---|
| Clock speed | up to 3200 MHz | from 4800 MHz |
| Bandwidth | lower | higher |
| Latency | often lower | slightly higher |
| Price | cheaper | more expensive |
| Availability | widespread | current platforms |
DDR5 offers more bandwidth, which matters for CPU-based AI workloads. DDR4 is sufficient for many entry-level setups.
Clock speed and timings
- Higher clock speeds deliver more bandwidth.
- Lower timings reduce latency.
- For AI work, clock speed typically matters more than extreme low-latency configurations.
- 3200 MHz DDR4 or 5600 MHz DDR5 are good starting points.
Channels
- Dual-Channel is the standard and important.
- Quad-Channel on workstation CPUs further improves bandwidth.
- Always populate matching module pairs.
ECC
ECC corrects memory errors and makes sense on workstation and server platforms. For pure AI inference in a homelab, it’s not strictly necessary, but it’s recommended for extended training runs or 24/7 operation.
UDIMM, RDIMM, LRDIMM
- UDIMM: Standard desktop modules.
- RDIMM: Registered modules for servers.
- LRDIMM: Reduced loading capacity for larger configurations.
Consumer CPUs use UDIMM, while workstation and server CPUs use RDIMM or LRDIMM.
RAM requirements for Ollama
- 7B model with Q4: 8-16 GB system RAM.
- 13B model: 16-32 GB.
- 70B model: 64-128 GB or more.
- Larger context windows and concurrent models increase requirements.
Tips
- Err on the side of more RAM rather than less.
- Always use dual-channel mode.
- Choose DDR5 if you have a newer platform.
- Use ECC on workstation platforms.
- Avoid populating memory with many small modules.
- With multi-GPU setups, allocate extra RAM for the OS and buffers.
- Verify that RAM clock speed is set correctly in the BIOS.
Common mistakes
- Single RAM module: Single-channel slows the entire system.
- Insufficient RAM: Models get swapped to disk.
- Wrong type: UDIMM won’t fit in an RDIMM slot.
- Incompatible clock speeds: System fails to boot or throttles.
- 2x16 instead of 4x8: Better for future expansion.
- XMP/EXPO not enabled: RAM runs slower than it should.
Further reading and resources
- BotServ.de CPU for AI
- BotServ.de GPU for local AI
- BotServ.de Building an AI PC
- BotServ.de AI Workstations
FAQ: RAM for AI
Do I need DDR5 for AI? Not strictly necessary, but the extra bandwidth helps.
Is ECC important? Yes for workstations and long-running jobs, less so for gaming homelabs.
How much RAM for 7B models? At least 16 GB if using a GPU, otherwise 32 GB.
Is 32 GB enough for Ollama? Yes for 7B and many 13B models. 70B requires significantly more.
Does faster RAM help? Yes, especially for CPU-based inference.
Sources and further reading
- DDR5: https://www.jedec.org/standards-documents/focus/memory/ddr5
- Crucial RAM Guide: https://www.crucial.com/articles/about-memory/speeds-and-feeds
Summary: Buying RAM for local AI
RAM is essential for fast, stable AI workloads. 7B models work fine with 16-32 GB, while larger or CPU-focused setups need considerably more. DDR5 and higher clock speeds benefit CPU workloads, while ECC provides stability on workstation platforms. By using dual- or quad-channel configurations, sizing appropriately, and checking compatibility, you’ll avoid memory bottlenecks and get the most from your hardware.


