Sizing Hardware Correctly: Planning Your AI PC
What this article covers
- How to plan hardware for local AI from the ground up, starting with your target model
- A formula to calculate the VRAM requirements of any model
- How GPU, RAM, CPU, SSD, and power supply interconnect and influence each other
- Three worked examples for budgets under €500, €500-1500, and €1500-3000
- A checklist and common pitfalls to help you avoid mistakes when purchasing
Introduction: Understanding hardware sizing
If you’re interested in running local AI, you’ll quickly face a deceptively simple question: what computer do I actually need? The answer is less obvious than it sounds. A popular gaming PC isn’t automatically a good AI PC, and an expensive workstation might be overkill for your purposes. The right approach doesn’t start with picking a graphics card. It starts with a different question entirely: which model do you want to run?
As you work through AI hardware fundamentals, you’ll encounter terms like VRAM, memory bandwidth, and GPU offloading. Each describes a piece of the puzzle in hardware planning. This article connects all the pieces and walks you through planning an AI PC that matches your needs, nothing more and nothing less.
You don’t need prior knowledge. We’ll work through each step together, from choosing your target model through VRAM calculation to selecting a power supply. By the end, you’ll have a concrete plan to execute.
Why does proper sizing matter?
Picture this: you buy a €1500 computer. Strong CPU, 32 GB of RAM, a nice graphics card with 8 GB of VRAM. The seller promises it’s perfect for AI. You install Ollama, load a 13-billion-parameter model, and fire it up. Then you wait. Every response takes several seconds, the model stutters, and you wonder why the expensive computer isn’t performing better. The issue: 8 GB VRAM isn’t enough for a 13B model, so it falls back to the CPU.
The opposite extreme: you spend €3000, buy the largest graphics card you can find, and end up only running 7B models that would work fine on a €300 card. You wasted €2200 because you didn’t think through your sizing.
Both scenarios stem from the same root cause. The hardware wasn’t matched to the target model. Proper sizing saves money, prevents frustration, and ensures your machine does exactly what you expect. If you already understand RAM vs. VRAM and CPU vs. GPU, you have the foundation. This article shows you how to apply that knowledge.
Hardware planning in brief
Planning hardware for local AI means selecting each component to match your target model. You start with the model and work downward. The model determines VRAM needs. VRAM needs determine the GPU. The GPU determines power consumption. Power consumption determines the power supply. Meanwhile, RAM, CPU, and SSD must be sized so they don’t become bottlenecks.
Think of it this way: imagine building a house. Before you order bricks, you need to know how many people will live there. A family of four needs a different house than a single person. If you buy bricks first and then decide who moves in, you’ll either build too small or too large. Hardware planning works the same way. Your target model is the family, and the components are the bricks. You plan from the use case, not from the parts.
Who should read this article?
This article is for beginners planning or buying a machine for local AI. If you’re wondering which graphics card you need, how much RAM is enough, or whether your current CPU even matters, you’re in the right place. You don’t need technical background knowledge, but it helps to have a rough grasp of the fundamentals in AI hardware basics.
If you’ve already experimented with Ollama and are thinking about an upgrade, this article helps. The worked examples and the checklist at the end give you concrete reference points so you’re not buying blind.
Key terms in hardware planning
| Term | Meaning |
|---|---|
| VRAM | Memory on the graphics card, exclusive to the GPU, determines model size |
| RAM | System memory used by the CPU and all programs |
| GPU | Processor on the graphics card with many parallel cores, ideal for AI |
| CPU | Main processor with fewer but more flexible cores |
| SSD | Fast storage where models reside |
| NVMe | Particularly fast SSD connected directly via PCIe |
| TDP | Thermal Design Power, the heat output of a component |
| Watt | Unit of power, here for system power consumption |
| PSU | Power Supply Unit, the supply that powers the machine |
| Form Factor | Physical form of a component, e.g., ATX for motherboards or dual-slot for graphics cards |
Step 1: Choose your target model
The first step is the most important and most often skipped. Before you pick any components, you must know which model you want to run. The model determines everything else, from VRAM through GPU to power supply.
Start with a simple question: what do you want the AI to do? Help with writing, generate code, or summarize documents? For simple text tasks, a 7B model like Llama 3.1 8B or Mistral 7B usually suffices. For more demanding work, like summarizing longer documents or answering complex questions, a 13B or 14B model makes sense. For very demanding tasks, such as multilingual translation or deep reasoning, consider a 70B model.
If you’re unsure, start small. A 7B model runs on almost every modern graphics card and gives you a feel for what local AI can do. You can always switch to a larger model later without changing hardware, as long as you leave some headroom in your planning. Learn more about the relationships at Model size and memory requirements.
Step 2: Calculate VRAM requirements
Once you know your target model, calculate its VRAM needs. The rule of thumb for quantized models, as typically loaded with Ollama, is:
VRAM needed in GB = Parameter count in billions × 0.6 to 0.8
The factor depends on quantization. For 4-bit quantization, estimate about 0.6 GB per billion parameters. For 5-bit, about 0.7 GB. For 6-bit, about 0.8 GB. Learn more about quantization in Quantization.
Add another 1 to 2 GB for context, the history of your conversation. Longer contexts require more memory.
Examples:
- 7B model, 4-bit: 7 × 0.6 = 4.2 GB, plus 1 GB context, about 5 GB VRAM
- 13B model, 4-bit: 13 × 0.6 = 7.8 GB, plus 1 GB context, about 9 GB VRAM
- 70B model, 4-bit: 70 × 0.6 = 42 GB, plus 2 GB context, about 44 GB VRAM
Add a safety margin of 10 to 20 percent so the model runs stably even during longer conversations. For precise numbers, see RAM and VRAM requirements.
Step 3: Choose Your GPU
The GPU is the core of your AI PC. It determines how fast your models run and how large they can be. VRAM is the critical metric, without enough of it, even the fastest GPU won’t help. Once you’ve settled on VRAM, then look at compute power and memory bandwidth.
Nvidia: Nvidia is the straightforward choice because most tools and frameworks are optimized for it. CUDA is the de facto standard in AI. Consumer cards like the RTX 4060 with 8 GB, the RTX 4070 with 12 GB, or the RTX 4090 with 24 GB cover most use cases. The downside is price, especially for cards with plenty of VRAM.
AMD: AMD GPUs work too, though they sometimes need more configuration. Tools like ROCm have improved things, but Nvidia remains the less complicated option. If you already own an AMD card, it’s worth trying. If you’re buying new and AI is your focus, Nvidia is usually the safer bet.
Apple Silicon: Macs with M-series chips use Unified Memory, where CPU and GPU share a single memory pool. A Mac Studio with 64 or 128 GB Unified Memory can load models that won’t fit on any consumer graphics card. That’s a strong option if you want to run large models without buying multiple GPUs. Learn more at Unified Memory.
Pick a GPU where your calculated VRAM requirement plus a safety margin fits comfortably. If your 13B model needs 9 GB, choose a card with 12 GB VRAM. If your 7B model needs 5 GB, 8 GB VRAM is enough. For more on selection, see the buying guide.
Step 4: Dimension Your RAM
RAM is your system’s working memory. It plays a minor role in inference speed as long as your entire model fits in VRAM. It becomes important when the model doesn’t fit in VRAM and parts spill to the CPU, that’s GPU offloading. Learn more in GPU Offloading.
A good rule of thumb: plan for at least double your VRAM requirement as RAM. If your model needs 9 GB VRAM, plan for at least 18 GB RAM, rounded up to 32 GB. This gives your system enough space to keep the model fully in RAM if the GPU isn’t sufficient, while still running normally.
DDR4 vs. DDR5: DDR5 is faster than DDR4 and offers higher memory bandwidth. This helps especially when parts of your model run on the CPU, since the CPU then accesses data faster. If you’re building a new system anyway, go DDR5. If you’re upgrading an existing machine, check whether your motherboard supports it. DDR4 works, but DDR5 is the more future-proof choice. For more on this connection, see Memory Bandwidth.
Step 5: CPU and Motherboard
The CPU plays a minor role in inference as long as your model runs on the GPU. It becomes important if parts of the model fall back to the CPU, and it determines how fast the model loads from RAM into VRAM. A modern CPU with enough cores keeps the whole system responsive.
Choose a CPU from a recent generation, such as an AMD Ryzen 5 or 7, or an Intel Core i5 or i7. You don’t need a flagship, but avoid very old or very weak CPUs. What matters is that the CPU provides enough PCIe lanes so the GPU can communicate at full speed. Most modern CPUs do this, but very budget models may have limitations.
The motherboard must match your CPU, with the right socket and enough PCIe slots. If you plan to add a second graphics card later, look for two x16 slots, though the second card often runs at reduced speed. Also check whether your motherboard supports DDR5 if you’re planning DDR5 RAM.
Step 6: SSD and Storage
Your SSD is where your models live. Each model is a file that gets loaded into RAM at startup and then into VRAM. The faster your SSD, the shorter the load time. A slow hard drive won’t affect inference speed itself, but it will be noticeable every time you load a model.
Go for an NVMe SSD connected directly via PCIe. NVMe SSDs are much faster than SATA SSDs and noticeably cut load times. A typical NVMe SSD reads at 3000 to 7000 MB/s, while a SATA SSD manages only around 500 MB/s.
How much storage do you need? A quantized 7B model is roughly 4 to 5 GB, a 13B model around 8 to 9 GB, a 70B model about 40 GB. If you want to store multiple models, plan for at least 1 TB. Models take space quickly, and you’ll likely experiment with more than you initially planned. 2 TB is a comfortable size that gives you room to grow.
Step 7: Power Supply and Cooling
Your power supply delivers electricity to your machine. Too weak and it will crash under load, especially during long inference runs. Too oversized and you’re wasting money. Calculate the right size from your components’ power consumption.
Add up the TDP values of all your parts. A typical setup looks like this:
GPU: 200 W
CPU: 100 W
RAM, motherboard, SSD: 50 W
Total: 350 W
Add a safety buffer of 30 to 50 percent so your PSU doesn’t run at the limit continuously. For 350 W total consumption, choose a 500 to 600 W power supply. For an RTX 4090 with 450 W TDP, you need at least 850 W, ideally 1000 W.
Pay attention to efficiency rating. A power supply with 80-Plus Gold certification is more efficient and runs cooler than a budget unit without certification. This saves power and extends lifespan.
Cooling keeps your components in the safe temperature range. Your GPU usually has its own cooling, but the whole machine needs adequate case fans to move hot air out. Look for a case with good airflow and at least two or three fans. If the GPU gets too hot, it throttles and slows down, which directly hurts inference speed.
Budget Examples for 3 Price Tiers
Here are three example builds to give you a sense of scale. Prices are approximate and will vary.
Under 500 euros: At this price point, you won’t build a dedicated AI PC, but rather use what you have or buy secondhand. A used Nvidia RTX 3060 with 12 GB VRAM costs around 200 euros. Add a simple machine with 16 GB RAM and an NVMe SSD. You’ll run 7B models smoothly and 13B models with GPU offloading. Another option is a Mac Mini M4 with 16 GB Unified Memory, starting at around 700 euros, just above budget, but very compact. For more on entry-level options, see AI PC Beginner.
500 to 1500 euros: At this level you get a solid AI PC. An Nvidia RTX 4060 with 8 GB VRAM costs around 300 euros, an RTX 4070 with 12 GB VRAM around 550 euros. Add a current-gen CPU, 32 GB DDR5 RAM, a 1 TB NVMe SSD, and a 650 W power supply. You’ll run 7B models smoothly and fit 13B models fully in VRAM if you go for the 12 GB card. A complete setup like this runs 1000 to 1300 euros.
1500 to 3000 euros: Here you’re considering an RTX 4090 with 24 GB VRAM, which costs around 2000 euros. You’ll run 13B and 14B models without issue, and 30B models fit in VRAM. For 70B models you need multi-GPU or Apple Silicon. Alternatively, buy two RTX 4070s with 12 GB each for 24 GB total, though this adds complexity and power draw. A Mac Studio with 64 GB Unified Memory starts around 2400 euros and is often the better choice for large models.
Hardware Checklist for AI
- Target model selected, including parameter count and quantization level
- VRAM requirements calculated, including context and safety margin
- GPU chosen with VRAM capacity matching the calculated need
- RAM at least double the VRAM requirement, rounded up to standard sizes
- CPU from current generation with sufficient PCIe lanes for the GPU
- Motherboard with compatible socket and DDR5 support if planned
- NVMe SSD with at least 1 TB, preferably 2 TB
- Power supply with 30 to 50 percent safety margin above calculated consumption
- Case with good airflow and at least two to three fans
- Power supply efficiency rating checked, minimum 80-Plus-Gold
- Total budget includes buffer for cables, cooling, and unexpected costs
Common Pitfalls in Hardware Planning
- Starting with the GPU instead of the model: Choosing a graphics card first and then figuring out what to run on it is backwards planning. The GPU should follow from the model, not the other way around.
- Underestimating VRAM requirements: Context consumes additional memory, and it grows during longer conversations. Calculating only the model size risks crashes during extended sessions.
- Undersizing RAM: When VRAM runs short, the CPU needs enough RAM to hold the remaining model layers. Insufficient RAM prevents the model from launching at all.
- Choosing an underpowered power supply: A weak PSU causes crashes under load. This is particularly frustrating because failures occur erratically and are hard to diagnose.
- Neglecting cooling: A hot GPU throttles and slows down. You won’t notice immediately, but inference speed drops noticeably.
- Using a slow SSD: A SATA SSD or HDD makes loading large models painfully slow. NVMe is essential for an AI PC.
- Ignoring quantization: Loading models without quantization requires significantly more VRAM. With 4-bit quantization, much larger models fit into VRAM with barely noticeable quality loss.
- Expecting VRAM upgrades: VRAM cannot be upgraded. If you need more VRAM, you must buy a new graphics card. Plan ahead with some headroom.
Hardware, Costs, and Privacy in Dimensioning
Hardware: Proper dimensioning starts with your target model and ends with the power supply. Every component must match the model, or you either create a bottleneck or waste money. The GPU is the most important component, but RAM, SSD, and power supply matter too. A strong AI PC is only as good as its weakest link.
Costs: The GPU is the biggest expense. An RTX 4060 with 8 GB costs roughly 300 euros, an RTX 4090 with 24 GB around 2000 euros. RAM is cheap, typically 3 to 5 euros per GB. NVMe SSDs cost about 50 to 80 euros per TB. Don’t skimp on the power supply, a quality 650W unit runs 80 to 120 euros. Overall, a solid AI PC costs 1000 to 1500 euros, while a high-end setup runs 2500 to 3000 euros.
Privacy: With locally running models, your data stays on your machine. No request goes to a server, and nothing leaves your system. This is a major advantage over cloud-based AI. Dimensioning plays no direct role here, the only requirement is that the model runs locally without calling external APIs.
Further Reading and Resources for Hardware Planning
- AI Hardware Basics for foundational knowledge about hardware for local AI
- RAM vs. VRAM to understand the difference between these two memory types
- CPU vs. GPU comparing the two processor types
- GPU Offloading to understand what happens when VRAM is insufficient
- Memory Bandwidth for the relationship between speed and storage
- Unified Memory explaining how Apple Silicon works
- Model Size and Memory Requirements for specific numbers on model sizes
- RAM and VRAM Requirements for concrete requirements per model
- Quantization to learn how models are compressed
- Buying Guide when planning a new graphics card or system
- Entry-Level AI PC for specific recommendations to get started
- Ollama the most popular tool for running models locally
FAQ: Dimensioning Hardware Correctly - Common Questions
How do I start hardware planning?
The first step is always the target model. Think about what you want to do with AI, pick a model, and calculate VRAM requirements. Only then look at graphics cards. Backwards planning, where you buy a GPU first and hunt for software later, usually wastes money.
How do I calculate a model’s VRAM requirement?
For quantized models, use this rule of thumb: parameter count in billions times 0.6 to 0.8 GB, depending on quantization. Add 1 to 2 GB for context and a 10 to 20 percent safety margin. A 7B model with 4-bit quantization needs about 5 GB, a 13B model around 9 GB.
How much RAM do I need for local AI?
Plan for at least double your VRAM requirement in RAM. If your model needs 9 GB of VRAM, choose at least 18 GB RAM, rounded up to 32 GB. This gives the system enough space to keep the model fully in RAM if the GPU falls short.
Is a CPU without a dedicated graphics card enough?
A CPU with sufficient RAM works for small 7B models but slowly. For serious local AI, a dedicated graphics card or Apple Silicon makes sense. GPUs are much faster because they have thousands of parallel cores.
Is DDR5 RAM worth it for local AI?
Yes, especially if the model partly runs on the CPU. DDR5 offers higher memory bandwidth than DDR4, speeding up CPU-based inference. If building a new system, choose DDR5. On an existing PC, check if your motherboard supports it.
How powerful does the power supply need to be?
Add up the TDP values of all components and add 30 to 50 percent safety margin. For a typical setup with 350W total consumption, choose a 500 to 600W power supply. For an RTX 4090, you need at least 850W, preferably 1000W.
Do I need an NVMe SSD or is SATA enough?
An NVMe SSD is recommended because it significantly speeds up loading large models. A SATA SSD works but is about six times slower. With a 40 GB model, that makes a noticeable difference at startup.
Can I upgrade later?
RAM and SSD can be upgraded easily. VRAM cannot, it’s soldered permanently to the graphics card. If you need more VRAM, you must buy a new GPU. Plan ahead with some headroom.
Is Apple Silicon worth it for local AI?
Yes, especially if you want to run large models. Apple Silicon uses Unified Memory, where the GPU can access all system RAM. A Mac Studio with 64 or 128 GB Unified Memory can load models that won’t fit on any consumer graphics card.
How much storage space do I need for models?
A quantized 7B model is about 4 to 5 GB, a 13B model roughly 8 to 9 GB, a 70B model around 40 GB. Plan for at least 1 TB, preferably 2 TB, so you can store multiple models without constantly freeing space.
What happens if the power supply is too weak?
An underpowered PSU causes crashes under load, especially during long inference runs. This is hard to diagnose because failures happen unpredictably. Always choose a power supply with a safety margin and check for good efficiency ratings.
Should I buy an Nvidia or AMD GPU?
Nvidia is the easier choice because most tools and frameworks are optimized for it. AMD GPUs work too but sometimes require extra configuration. If you’re buying new and AI is the focus, Nvidia is usually the safer bet.
References and Further Reading
- Nvidia specifications for RTX 4060, 4070, and 4090
- AMD Ryzen specifications and PCIe lanes
- Apple Silicon Unified Memory overview
- Ollama documentation on model loading and GPU offloading
- Insights from the local AI community on hardware configurations


