PC Hardware for Ollama: Optimizing for Local Models
What this article covers
- Which hardware components actually matter for smooth Ollama inference
- How much VRAM and RAM you need for 3B, 7B, 13B, 30B, and 70B models
- Which GPUs are worth the investment for Ollama, from RTX 3060 to RTX 4090
- How Apple Silicon performs as an Ollama platform
- Three concrete builds with pricing for every budget
Introduction: Understanding PC Hardware for Ollama
Ollama has made running language models locally dramatically simpler. One command, and a model downloads and starts. The real question is what hardware to run it on. Start Ollama on a regular laptop without a GPU, and you’ll wait seconds between words. With the right graphics card, those same models respond as fluidly as a typical chatbot.
This article explains what kind of PC makes sense for Ollama, how much VRAM and RAM you need for each model size, and which builds suit different budgets. For a broader look at AI hardware, check out our buying guide.
Why does Ollama need specialized hardware?
Ollama runs on any modern computer. That doesn’t mean it runs fast. Running a 7B model like Llama 3.1 on CPU alone typically yields 5 to 15 tokens per second. It feels like watching someone hunt and peck.
The same model on an RTX 3060 with 12 GB VRAM delivers 40 to 80 tokens per second. That’s a difference you notice immediately. The GPU handles computation for model layers while the CPU recedes into the background. Once the entire model fits in VRAM, inference becomes genuinely snappy.
The reason is straightforward: LLM inference is memory bandwidth limited. A GPU like the RTX 4090 offers over 1,000 GB/s of memory bandwidth, while typical DDR5 workstation RAM maxes out around 80 to 100 GB/s. More bandwidth means more tokens per second. For deeper context, see our article on CPU vs. GPU.
Ollama hardware requirements at a glance
For Ollama, two things matter most: enough storage for the model and sufficient bandwidth to process it quickly. In practice: a GPU with 12 GB VRAM runs 7B and 8B models in Q4-K quantization smoothly. 24 GB VRAM handles 13B models and some 30B variants. If you want to run 70B models locally, you’ll need 48 GB VRAM or more, or alternatively a Mac with 64 GB Unified Memory or higher.
RAM is the second critical factor. If the model doesn’t fit entirely in VRAM, Ollama offloads pieces to CPU. This works but gets slower. As a rule of thumb: have at least as much RAM as VRAM, ideally double. Our article on RAM and VRAM requirements goes deeper.
Who should read this article
This article targets beginners and intermediate users planning or upgrading a PC for Ollama. You don’t need to be an AI expert. If you understand what RAM and VRAM are, that’s enough. For those completely new to the topic, our guide AI PC for Beginners provides broader context.
Key terminology for Ollama hardware
| Term | Definition |
|---|---|
| VRAM | GPU video memory. Determines how much of the model lives on your graphics card. |
| RAM | System memory. Used when the model doesn’t fit entirely in VRAM. |
| GPU | Graphics card that accelerates inference. For Ollama, usually NVIDIA with CUDA. |
| CPU | Main processor. Runs model layers that don’t fit on the GPU. |
| Quantization | Reduces model precision to save memory. See Quantization for details. |
| GPU Offloading | Moving model layers from CPU to GPU. Explained in GPU Offloading. |
| OLLAMA_GPU_LAYERS | Environment variable controlling how many layers shift to the GPU. |
| KV-Cache | Stores precomputed attention values during inference. Grows with context length. |
| Token/s | Speed metric: how many tokens generated per second. |
| NVMe | Fast SSD storage. Important for quick load times when starting large models. |
Ollama hardware requirements by model size
The table below shows typical needs for common model sizes in Q4-K quantization. These are guidelines; actual requirements depend on the specific model and context length.
| Model Size | VRAM (Q4) | RAM (Minimum) | Recommended GPU |
|---|---|---|---|
| 3B | 3 to 4 GB | 8 GB | RTX 3050, integrated GPU |
| 7B | 5 to 6 GB | 16 GB | RTX 3060 12 GB |
| 8B | 6 to 7 GB | 16 GB | RTX 3060 12 GB, RTX 4060 |
| 13B | 9 to 10 GB | 32 GB | RTX 3060 12 GB, RTX 4070 |
| 30B | 20 to 22 GB | 32 GB | RTX 3090 24 GB, RTX 4090 |
| 70B | 40 to 48 GB | 64 GB | 2x RTX 3090, Mac with 64 GB+ |
For a deeper dive into requirements, see Ollama Hardware Requirements.
Ollama on CPU: What’s possible
Ollama runs on CPU, but it’s slow. Small 3B models work reasonably on a modern processor, delivering around 10 to 20 tokens per second. 7B models achieve 5 to 15 tokens per second depending on CPU and RAM speed. 13B and larger become uncomfortable, with several seconds of wait between sentences.
One caveat: even on CPU, faster RAM helps. DDR5 at high clock speeds noticeably outpaces older DDR4. Without a GPU, plan for at least 32 GB RAM, preferably 64 GB, so larger models can load at all.
Ollama with GPU: Which graphics card?
A dedicated GPU is the biggest lever for Ollama performance. NVIDIA cards with CUDA are the standard because most models are optimized for them.
- RTX 3060 12 GB: The entry point for price and performance. 12 GB VRAM handles 7B, 8B, and 13B in Q4. Often affordable secondhand.
- RTX 4070 12 GB: Faster than the 3060, same VRAM. Good if you want more performance per token.
- RTX 4090 24 GB: High-end for enthusiasts. 24 GB VRAM fits 30B models fully, 70B only with CPU offloading.
- Apple Silicon: Not a traditional GPU, but Unified Memory makes Macs a compelling Ollama platform. More below.
Combining multiple GPUs lets you add VRAM. Two RTX 3090s with 24 GB each yield 48 GB VRAM, enough for 70B models in Q4. This is a popular used-market configuration.
Ollama on Apple Silicon
Apple Silicon is a compelling platform for Ollama. The Unified Memory approach means CPU and GPU share the same memory pool. A Mac with 64 GB Unified Memory can dedicate much of it as VRAM. This isn’t possible with traditional PCs using dedicated GPUs.
- Mac mini M4 with 24 GB: Good for 7B and 8B, tight for 13B.
- Mac Studio M2 Ultra with 128 GB: Runs 70B models in Q4 smoothly, one of the most compact solutions for large models.
- MacBook Pro M4 Pro with 48 GB: Mobile and powerful, suitable for 13B and 30B.
The tradeoff: Apple Silicon is expensive and not upgradeable. Want more memory later? You need a new machine. In return, Macs are quiet, efficient, and compact.
Recommended Hardware from Amazon Shop
Hardware for Ollama on Amazon Shop
Bei Amazon ansehenAffiliate-Link: Bei einem Kauf erhalten wir möglicherweise eine Provision.
Build Suggestions for Ollama
Build 1: Budget, around EUR 600 to 800
- CPU: AMD Ryzen 5 5600 or Intel Core i5-12400F
- RAM: 32 GB DDR4
- GPU: RTX 3060 12 GB (used or new)
- Storage: 1 TB NVMe SSD
- Power supply: 550 W, 80+ Bronze
This build handles 7B and 8B models smoothly. 13B fits in VRAM with Q4 quantization, though snugly. Perfect for getting started and experimenting with your first AI agents.
Build 2: Mid-range, around EUR 1,200 to 1,800
- CPU: AMD Ryzen 5 7600 or Intel Core i5-13600K
- RAM: 64 GB DDR5
- GPU: RTX 4070 12 GB or used RTX 3090 24 GB
- Storage: 2 TB NVMe SSD
- Power supply: 750 W, 80+ Gold
With 24 GB VRAM (RTX 3090), 13B and 30B models run comfortably. 64 GB RAM gives you room for offloading and parallel processes. A solid all-rounder for serious local AI work.
Build 3: High-end, around EUR 3,000 to 5,000
- CPU: AMD Ryzen 9 7950X or Intel Core i9-14900K
- RAM: 128 GB DDR5
- GPU: 2x RTX 4090 24 GB or 2x RTX 3090 24 GB
- Storage: 4 TB NVMe SSD
- Power supply: 1,200 W, 80+ Platinum
48 GB VRAM lets you run 70B models in Q4. 128 GB RAM ensures even long contexts and multiple concurrent models pose no problem. If you want to squeeze every bit of performance from Ollama, this is the setup for you.
Ollama Performance Tips
- Maximize GPU layers: Set
OLLAMA_GPU_LAYERSto a high value or-1so Ollama offloads as many layers as possible to the GPU. Configuration details are at Ollama Configuration. - Choose your quantization wisely: Q4_K_M is a solid default. Q8 offers more accuracy but demands significantly more VRAM. Q3 saves space at the cost of quality.
- Limit context length: The KV-cache grows with context size. A 32k token context can consume several extra GB of VRAM. Reduce context if you need speed.
- Use NVMe over SATA SSD: Large models load much faster from NVMe. You’ll notice the difference at startup and when switching between models.
- Close other applications: Browsers with many tabs and other software compete for RAM and VRAM. Close them before loading large models.
Common Hardware Pitfalls with Ollama
- Underestimating VRAM needs: If you have exactly 8 GB VRAM and want to load an 8B model in Q8, you’re out of luck. Q4 fits, Q8 doesn’t. Always budget extra headroom for the KV-cache.
- Overlooking RAM requirements: When a model doesn’t fit in VRAM, Ollama falls back to RAM. With 16 GB RAM and a 12 GB GPU, you’ll hit limits quickly with 13B models.
- Underestimating CPU-offloading impact: When only part of the layers fit on the GPU, inference slows down dramatically. A 30B model split between CPU and GPU becomes frustratingly slow.
- Choosing an undersized power supply: Two RTX 3090 cards draw over 700 W under load alone. A weak PSU leads to crashes or unexpected shutdowns.
- Case too small: Large GPUs don’t fit in every enclosure. Measure beforehand, especially for dual-GPU setups.
- Neglecting cooling: Two GPUs in a tight case get hot. Fan control and airflow matter, otherwise the card will throttle.
- Apple Silicon memory is fixed: Buy a Mac with too little memory and you’re stuck forever. Think ahead about what model sizes you want to run.
- AMD GPUs with ROCm: Ollama supports AMD, but CUDA is better optimized. If compatibility matters most, go NVIDIA.
Hardware, Cost, and Security with Ollama
Running local inference with Ollama means your data stays on your machine. That’s a major advantage over cloud services. Chats, prompts, and contexts live on your disk.
From a cost perspective, hardware is a one-time investment with no API fees afterward. A EUR 800 build pays for itself quickly if you work regularly with large models. Factor in electricity costs: an RTX 4090 draws up to 450 W under load, and two of them pull significantly more.
On security: models from unknown sources can contain malicious code. Download models only from trusted repositories like the official Ollama library or Hugging Face from verified uploaders.
Further Reading and Ollama Hardware Resources
- Ollama - Basics and setup
- Ollama Hardware Requirements - Detailed specifications
- Ollama Configuration - Settings and tuning
- Hardware Buying Guide - More hardware recommendations
- AI PC for Beginners - Broader introduction to AI hardware
- CPU vs. GPU - Why GPU makes the difference
- GPU Offloading - Layer distribution explained
- Quantization - How models get smaller
- RAM and VRAM Requirements - Memory needs in detail
FAQ: PC for Ollama - Common Questions
What GPU do I need for Ollama?
For 7B and 8B models, an RTX 3060 with 12 GB VRAM suffices. For 13B, 12 to 24 GB VRAM is recommended. 30B models need at least 24 GB. 70B models require 48 GB VRAM or more.
Can I use Ollama without a GPU?
Yes, Ollama runs on CPU. Performance is significantly slower. 7B models reach 5 to 15 tokens per second, and larger models become uncomfortable.
How much RAM do I need for Ollama?
Minimum 16 GB for small models on CPU. For 7B to 13B, 32 GB is recommended. For 30B and 70B, 64 GB or more. If the model doesn’t fit in VRAM, Ollama uses RAM as overflow.
Is an RTX 3060 enough for Ollama?
Yes, the RTX 3060 with 12 GB VRAM is an excellent entry point. It runs 7B, 8B, and 13B in Q4_K quantization smoothly. For 30B, it gets tight, and CPU offloading helps.
Is an RTX 4090 worth it for Ollama?
For enthusiasts, yes. 24 GB VRAM fits 30B models entirely and 70B with offloading. If you want maximum performance, the 4090 is the right choice. For beginners, it’s overkill.
Is Apple Silicon good for Ollama?
Yes, especially thanks to Unified Memory. A Mac Studio with 128 GB can run 70B models, which is hard to achieve with a single graphics card. The downside is price and no upgrade path.
What is GPU offloading in Ollama?
GPU offloading moves model layers from CPU to GPU. The more layers that fit on the GPU, the faster inference becomes. The OLLAMA_GPU_LAYERS environment variable controls how many layers get offloaded.
Can I use multiple GPUs for Ollama?
Yes, Ollama supports multiple GPUs. VRAM stacks. Two RTX 3090 cards with 24 GB each give 48 GB VRAM, enough for 70B models in Q4. Watch your power supply and cooling.
What is quantization in Ollama?
Quantization reduces the precision of model weights to save memory. Q4_K_M is a good balance between size and quality. Q8 is more accurate but uses roughly twice the VRAM as Q4.
How fast is Ollama on CPU?
On a modern CPU, you get about 5 to 15 tokens per second with 7B models. 3B models reach 10 to 20 tokens per second. 13B and larger become slow, often below 5 tokens per second.
Do I need an NVMe SSD for Ollama?
Not strictly necessary, but recommended. Large models load much faster from NVMe than SATA SSD. Switching between models saves noticeable time.
Does Ollama work with AMD graphics cards?
Yes, Ollama supports AMD through ROCm. Compatibility is not as broad as NVIDIA with CUDA. If you want to play it safe, choose an NVIDIA GPU.
References and Further Reading
- Ollama Model Library
- Ollama Documentation
- NVIDIA GPU Specifications
- Apple Silicon Overview
- Hugging Face Models
- Hardware for Ollama on Amazon Shop


