Ollama Hardware Requirements: What You Actually Need
What this article covers
- VRAM and RAM requirements for different model sizes, from 3B to 70B parameters
- When a GPU makes sense and when CPU-only is sufficient
- Recommended setups for various budgets, from under 500 to over 2000 euros
- How to verify with simple commands whether your hardware can handle a specific model
- Common pitfalls that prevent models from running despite seemingly adequate hardware
Understanding Ollama hardware requirements
Ollama makes it straightforward to run local language models. You download a model, start it, and go. But between that first ollama run and a smooth response lies a set of hardware decisions that determine success or frustration. A 7B model on a laptop without a GPU can take minutes for a single answer. A 70B model on a graphics card with insufficient VRAM will crash or limp along with CPU fallback.
This article gives you concrete numbers, recommendations, and real-world values so you buy or use exactly the hardware you need. Nothing more, nothing less. If you haven’t installed Ollama yet, the Ollama Installation guide has the steps you need.
Why these requirements matter
Imagine buying an RTX 4060 with 8 GB VRAM because a forum recommended a graphics card for local AI. You install Ollama, download llama3:70b, and start the model. Nothing happens. VRAM fills up, Ollama spills parts into system RAM, and token generation drops to 1 or 2 tokens per second. You wait, give up, and wonder what went wrong.
The issue: nobody told you that a 70B model at 4-bit quantization needs roughly 40 GB of memory. Your 8 GB VRAM works for a 7B model, not 70B. Without knowing the requirements, you buy blind, and the model either won’t run or runs painfully slowly. This article prevents exactly that.
Ollama hardware requirements explained
Ollama needs three things: enough memory for the model, enough compute power for inference, and enough disk space for model files. The critical factor is memory. A model must fit in VRAM (on GPU) or system RAM (on CPU) to run at acceptable speed. The rule of thumb: a 4-bit quantized model needs about 0.7 GB of memory per billion parameters. A 7B model thus needs around 5 GB; a 13B model about 9 GB.
GPU offloading speeds up inference tremendously, but only works if the model or most of it fits in VRAM. Leftovers spill into system RAM and get computed by the CPU, which is much slower. For the theory behind this, see RAM and VRAM Requirements and Quantization.
Who this article is for
This article is for beginners trying out Ollama or expanding their setup. You don’t need to be a computer scientist to follow the tables and recommendations. If you want to dive deeper into the fundamentals, check out AI Hardware Basics.
Key terms for Ollama hardware
| Term | Meaning |
|---|---|
| VRAM | Video memory on the graphics card, critical for GPU inference |
| RAM | System memory, used for CPU inference or GPU offloading overflow |
| GPU | Graphics card that accelerates model computation |
| CPU | Main processor that computes the model when no GPU is available or VRAM runs out |
| SSD | Solid state drive, much faster than HDD for loading models |
| NVMe | Very fast SSD interface, recommended for large models |
| Quantization | Model compression to 4-bit or lower, significantly reduces memory requirements |
| GPU Offloading | Moving model layers to the GPU, speeds up inference |
| Unified Memory | Shared memory on Apple Silicon that combines RAM and VRAM on one chip |
| OLLAMA_GPU_LAYERS | Environment variable that controls how many model layers load to GPU |
Minimum requirements
You don’t need much for Ollama to run at all. The bare essentials are:
- CPU: A modern x86-64 processor or Apple Silicon, at least 4 cores
- RAM: 8 GB, preferably 16 GB to avoid hitting limits immediately
- Storage: 10 GB free space for Ollama itself and a small model
- GPU: Not required; Ollama runs fine on CPU alone
- OS: Linux, macOS, or Windows
With this setup you can run small models like qwen2.5:3b or phi3. Performance is moderate but functional. A 7B model will load on a CPU-only system with 8 GB RAM, but noticeably slow.
Recommended hardware by model size
The table below shows how much VRAM and RAM you need for different model sizes. Values are for 4-bit quantization, which Ollama uses by default.
| Model size | VRAM (GPU) | RAM (CPU-only) | Recommended GPU | Recommended RAM |
|---|---|---|---|---|
| 3B | 3 GB | 6 GB | Not required | 8 GB |
| 7B | 5 GB | 10 GB | RTX 3060 12 GB | 16 GB |
| 8B | 6 GB | 12 GB | RTX 3060 12 GB | 16 GB |
| 13B | 9 GB | 18 GB | RTX 4060 Ti 16 GB | 32 GB |
| 30B | 20 GB | 36 GB | RTX 4090 24 GB | 64 GB |
| 70B | 40 GB | 72 GB | 2x RTX 4090 | 128 GB |
These are guidelines. Actual requirements depend on the specific quantization and model architecture. Mistral-7B uses slightly less memory than Llama-3-8B, but the ballpark is correct. For more details on sizing, see Sizing Hardware Correctly.
Ollama on CPU: what’s possible?
Running Ollama on a CPU is entirely legitimate, especially for smaller models and testing. Speed depends heavily on the processor and model size. With a modern 6-core CPU and 32 GB RAM, you get roughly:
- 3B model: 15 to 25 tokens per second, smooth to use
- 7B model: 5 to 10 tokens per second, acceptable for chat
- 13B model: 2 to 4 tokens per second, requires patience
- 30B and larger: Below 1 token per second, practically unusable
If you only chat occasionally with a 7B model and lack a GPU, CPU-only works. For regular use or bigger models, a GPU pays off. See CPU vs. GPU for a deeper comparison.
Ollama with GPU: which graphics card?
A GPU multiplies inference speed. The key metric is VRAM, not raw compute power. An RTX 3060 with 12 GB VRAM often outperforms an RTX 4060 with 8 GB because it fits larger models entirely in VRAM.
RTX 3060 12 GB: Entry-level recommendation. Enough VRAM for 7B and 8B models with headroom. Around 300 euros used. Achieves 40 to 60 tokens per second on a 7B model.
RTX 4060 Ti 16 GB: Mid-range with generous VRAM. Fits 13B models entirely in VRAM. Around 450 euros. A good choice if you want to run more than just small models.
RTX 4090 24 GB: High-end for serious users. Handles 30B models fully and 70B models with partial CPU offloading. From 1700 euros. Exceeds 100 tokens per second on a 7B model.
Apple Silicon: See the next section, since Unified Memory changes the math here.
More recommendations by budget are in the Buying Guide and Beginner AI PC.
Ollama on Apple Silicon
Apple Silicon uses Unified Memory, which means RAM and VRAM are physically the same pool of memory. A Mac with 32 GB RAM makes nearly 32 GB available to models, minus space reserved by the operating system and other applications. This makes Macs particularly appealing for large models.
Mac mini M4 with 24 GB: Handles 7B and 8B models well, delivering smooth performance at 30 to 50 tokens per second. 13B models run at reduced speed. Starting at around 700 EUR.
Mac Studio M2 Max with 64 GB: Runs 30B models entirely within Unified Memory. 70B models are feasible with some offloading. Starting at around 2200 EUR.
MacBook Pro M4 Pro with 48 GB: Mobile and powerful. Works well with 13B and 30B models. Starting at around 2000 EUR.
The advantage of Apple Silicon is that you don’t need a separate GPU. The downside is cost, and the fact that Nvidia GPUs with comparable VRAM are often faster. For more on storage considerations, see RAM vs. VRAM.
Disk Space for Models
Models consume significant storage. Ollama stores them by default in ~/.ollama/models. A model file’s size roughly matches its runtime memory footprint:
- 3B model: 2 to 3 GB
- 7B model: 4 to 5 GB
- 13B model: 7 to 9 GB
- 30B model: 17 to 20 GB
- 70B model: 38 to 42 GB
When you install multiple models, space requirements add up quickly. An NVMe SSD is recommended, since loading a 70B model from a hard drive can take several minutes. With NVMe, it takes seconds. You can change the storage location using the OLLAMA_MODELS environment variable; see Ollama Configuration for details.
Recommended Setups by Budget
Under 500 EUR
An existing PC or laptop with 16 GB RAM and an NVMe SSD. No dedicated GPU. You run 3B and 7B models on the CPU. Performance is moderate but functional. Ideal for learning and experimentation.
500 to 1000 EUR
A PC with 32 GB RAM and a used RTX 3060 with 12 GB. This lets you run 7B and 8B models entirely on the GPU, achieving 40 to 60 tokens per second. 13B models work with partial CPU offloading.
1000 to 2000 EUR
A PC with 64 GB RAM and an RTX 4060 Ti with 16 GB, or a used RTX 4090. You run 13B models fully on the GPU and 30B models with some offloading. This setup covers most use cases.
Over 2000 EUR
A system with 128 GB RAM and two RTX 4090s, or a Mac Studio with 64 GB Unified Memory. You can run 70B models, though with partial offloading. This setup is for serious users who regularly work with large models.
How Do I Check if My Hardware Is Sufficient?
Before loading a model, you can check your hardware. On Linux and Windows with an Nvidia GPU:
nvidia-smi
This shows your VRAM usage and GPU utilization. After starting a model, check how much VRAM it occupies:
ollama ps
This lists all running models with their memory consumption. On Linux, free -h displays system RAM:
free -h
On macOS, open Activity Monitor or run in the terminal:
vm_stat
If VRAM is 100% full after loading a model but token generation is still slow, Ollama is offloading parts to system RAM. In this case, you need more VRAM or a smaller model.
Common Hardware Pitfalls
-
Underestimating VRAM: Many people buy a GPU with 8 GB VRAM and wonder why 13B models run slowly. 8 GB is enough for 7B, not 13B.
-
Ignoring quantization: An unquantized 7B model needs about 14 GB; the same model in 4-bit needs only 5 GB. Without knowing the defaults, you can massively overestimate or underestimate requirements.
-
Insufficient system RAM: Even with CPU-only inference, system RAM must accommodate the operating system. 16 GB total RAM minus 9 GB for a 13B model leaves only 7 GB for everything else. This causes excessive swapping and severe slowdowns.
-
Slow storage: Loading models from a hard drive takes minutes with large models. While not a permanent problem, it’s frustrating when experimenting with different models.
-
Missing GPU drivers: Ollama detects your GPU only if drivers are properly installed. Without them, everything runs on the CPU, even with an expensive GPU in the system.
-
OLLAMA_GPU_LAYERS not set: By default, Ollama tries to load as many layers as possible to the GPU. For some models, manually limiting the layer count helps avoid crashes when VRAM fills up.
-
Forgetting Unified Memory on Apple Silicon: A Mac with 16 GB RAM doesn’t have 16 GB available for models. The operating system claims part of it, leaving realistically 10 to 12 GB.
-
Running multiple models simultaneously: When you run two models in parallel, they share VRAM. What fits individually may exceed capacity together.
Hardware, Cost, and Privacy with Ollama
Local AI means your data stays on your computer. No API calls to external servers, no cloud dependency. This is a privacy advantage, but it places hardware responsibility on you. A 2000-EUR system for 70B models makes sense only if you regularly need those models. For occasional use, a 500-EUR setup suffices.
Hardware costs are one-time expenses, unlike cloud APIs where you pay per token. With heavy usage, a local setup pays for itself quickly. The Ollama Overview gives you a full picture of available features.
Further Reading and Hardware Requirements
- Ollama Overview - All Ollama features at a glance
- Ollama Installation - How to install Ollama
- Ollama Configuration - Paths, ports, and environment variables
- RAM and VRAM Requirements - Storage fundamentals
- Quantization - How compression reduces memory needs
- AI Hardware Fundamentals - General basics
- CPU vs. GPU - When to use each option
- RAM vs. VRAM - Differences and similarities
- Sizing Hardware Correctly - Practical sizing
- Buying Guide - Recommendations by budget
- AI PC for Beginners - Your first AI PC
FAQ: Ollama Hardware Requirements - Common Questions
Can I use Ollama without a GPU?
Yes, Ollama runs entirely on the CPU. Small models up to 7B are usable at moderate speeds. Larger models become very slow.
How much VRAM do I need for a 7B model?
About 5 GB with 4-bit quantization. An RTX 3060 with 12 GB VRAM has enough headroom and is a solid choice.
Is 8 GB of RAM enough for Ollama?
For small models like 3B, yes. For 7B models it gets tight, especially when the operating system needs RAM too. 16 GB is the recommended minimum.
Does a single RTX 4090 run a 70B model?
Not entirely. A 70B model needs about 40 GB; the RTX 4090 has 24 GB VRAM. Ollama offloads the remainder to system RAM, which significantly reduces speed.
Is Apple Silicon better than Nvidia for Ollama?
With comparable memory, Nvidia is often faster. Apple Silicon has the advantage of Unified Memory, which enables very large models without buying multiple GPUs.
How much disk space do I need?
Budget 2 to 42 GB per model depending on size. For multiple models, plan for at least 100 GB of free space on an NVMe SSD.
What is GPU offloading?
GPU offloading means parts of the model load onto the GPU while the rest stays in system RAM. The more layers that fit on the GPU, the faster the model runs.
Do I need CUDA for Ollama?
On Linux with Nvidia GPUs, yes; CUDA drivers must be installed. Ollama detects them automatically. On macOS and for CPU-only operation, CUDA is not needed.
Can I run Ollama on a laptop?
Yes, if the laptop has enough RAM or VRAM. Laptops with dedicated GPUs like an RTX 4060 are well-suited. Without a GPU, small models work on the CPU.
What does OLLAMA_GPU_LAYERS mean?
This is an environment variable that controls how many model layers load onto the GPU. When VRAM is tight, reduce the value to avoid crashes.
How fast is Ollama on the CPU?
On a 7B model with a modern 6-core CPU, you’ll reach about 5 to 10 tokens per second. Usable for chat, but not as smooth as with a GPU.
Is a used GPU worth it?
Yes, especially a used RTX 3060 with 12 GB offers lots of VRAM for little money. Make sure the card works and drivers are up to date.
Sources and Further Reading
- Ollama official documentation
- Nvidia CUDA Toolkit documentation
- Apple Developer documentation on Unified Memory
- Benchmark results from our own measurements across various models and hardware configurations


