AI Hardware Fundamentals
What This Article Covers
- Which hardware components matter for local AI and why.
- How CPU, GPU, RAM, VRAM, and SSD work together.
- What TOPS, APU, and Unified Memory mean.
- Links to all fundamentals articles and buying guides.
Introduction
Running AI models locally requires the right hardware. A language model is a large file that gets loaded into memory and processed there. Whether it runs smoothly depends on your CPU, GPU, RAM, VRAM, and storage. This section gathers foundational knowledge to help you understand, assemble, or purchase an AI PC.
Why Learn AI Hardware Basics?
Without this knowledge, you’re buying blind. You risk ending up with a machine that’s either too weak for your models or too expensive for what you actually need. Take this example: running Llama 3.1 8B requires roughly 8 GB of memory for the quantized model. On a GPU with 8 GB VRAM, it runs smoothly. On a machine with only 16 GB RAM using CPU inference, each response takes several seconds. Knowing this upfront saves money and frustration.
These basics also help you evaluate marketing claims. A mini-PC with 128 GB RAM sounds impressive, but if memory bandwidth is low, a large model still runs slowly. TOPS values on a datasheet don’t tell the whole story because they measure peak performance under ideal conditions.
AI Hardware Explained Simply
Local AI loads a model into memory and computes responses from it. There are two approaches:
- CPU Inference: The model sits in regular system RAM. The CPU calculates the responses. Slower, but flexible and cheap.
- GPU Inference: The model sits in graphics memory (VRAM). The GPU calculates the responses in parallel, much faster.
Memory capacity determines which models fit at all. Compute power determines how fast you get answers. An SSD ensures the model loads quickly. An APU like Apple Silicon or Ryzen AI Max combines CPU and GPU on one chip and shares the same memory pool.
Who Should Read This?
- You want to run AI locally and understand what your machine can do.
- You’re considering buying a new PC or mini-PC.
- You’re new to this and want to make sense of marketing claims.
- You need no prior knowledge. Technical terms are explained as they come up.
Key Components
| Component | Role | What to Look For |
|---|---|---|
| CPU | Runs inference and operating system | Modern cores, plenty of cache |
| RAM | Stores models during CPU inference | 16 GB minimum, 32 GB better, 64-128 GB ideal |
| GPU/VRAM | Fast parallel computation | 8 GB minimum, 16 GB or more preferred |
| APU | CPU and GPU on one chip | Unified or Shared Memory uses common RAM |
| SSD | Fast storage for models | 500 GB or more, preferably NVMe |
Key Terms
- TOPS/NTOPS - Tera Operations Per Second. The unit of measurement for AI compute power. Higher values mean faster model execution.
- APU - Accelerated Processing Unit: CPU and GPU on one chip, like Apple Silicon or Ryzen AI Max.
- Unified/Shared Memory - RAM is used jointly by CPU and GPU. Essential for Apple Silicon and Ryzen AI Max.
- Quantization - Reduces model precision to run it with less memory. Lets larger models fit on smaller hardware.
- VRAM - Memory on a dedicated graphics card. Critical for GPU inference.
- LLM - Large Language Model.
Content and Articles
- RAM vs. VRAM - The difference between them, where models live, and why both matter for local AI.
- CPU vs. GPU - Which processor is faster for local AI and when CPU alone suffices.
- GPU Offloading - When VRAM runs short: splitting models between GPU and RAM.
- Memory Bandwidth - The most important factor in AI inference speed.
- Unified Memory - How Apple Silicon and Ryzen AI Max merge RAM and VRAM.
- Model Size and Memory Requirements - How large AI models really are and how much memory they need.
- Sizing Hardware Right - Step-by-step to the right AI PC for your needs.
- AI Fundamentals Glossary - All important terms A to Z explained clearly.
Related articles:
- RAM and VRAM Requirements - What your machine actually needs for local models.
- Quantization - How large models fit on smaller hardware.
- Beginner AI PC - Concrete recommendations and price tiers.
- Buying Guide Overview - What to consider when purchasing.
Common Pitfalls
- RAM is not VRAM - 32 GB RAM helps little if your GPU has only 4 GB VRAM. For GPU inference, graphics memory is what counts.
- TOPS don’t tell the whole story - A high TOPS value doesn’t guarantee your model runs fast. Memory bandwidth and architecture matter equally.
- Mini-PCs aren’t always upgradeable - If you want more memory later, you often need a new device. Choose enough RAM from the start.
- Apple Silicon is efficient, but not everything runs natively - Some open-source models are optimized for CUDA first and run better on Linux or Windows with an Nvidia GPU.
Further Reading
FAQ - Common Questions
What’s the difference between RAM and VRAM?
RAM is your machine’s main memory, VRAM is the graphics card’s memory. During CPU inference, the model sits in RAM. During GPU inference, it sits in VRAM. GPU inference is much faster but needs enough VRAM. Read more in RAM vs. VRAM.
What matters more: RAM or VRAM?
Both. CPU-based inference needs lots of RAM, GPU-based inference primarily needs VRAM. For larger models, a graphics card with sufficient VRAM is usually faster.
What does TOPS mean?
TOPS stands for Tera Operations Per Second. It measures how many compute operations a chip can perform per second for AI tasks. Higher values mean potentially faster AI inference, but memory bandwidth and architecture are equally important.
Do I need a graphics card for local AI?
Not necessarily. Smaller models like Llama 3.1 8B run on the CPU alone. For faster responses and larger models, a GPU with sufficient VRAM is recommended.
What is an APU?
An APU combines CPU and GPU on one chip. Examples include Apple Silicon and AMD Ryzen AI Max. The advantage is both processors can access the same memory, which helps significantly with local AI.
What is Unified Memory?
With Unified Memory, CPU and GPU share the same system RAM. This means a Mac with 32 GB RAM can use all 32 GB for the GPU as well. With a traditional graphics card, VRAM is separate and limited.
How much RAM do I need for local AI?
For 7B models, at least 16 GB RAM; 32 GB is better. Larger models like 70B need 64 GB or more. With cloud APIs, you only need memory for the framework itself.
How much VRAM do I need for local AI?
For quantized 7B models, at least 8 GB VRAM. For 13B models, 12 GB or more. 70B models need 24 GB or more, often multiple GPUs.
Why is an SSD important?
A language model can be several gigabytes in size. An NVMe SSD loads the model much faster than an HDD. This reduces startup time and keeps your workflow smooth.
Is a mini-PC worth it for local AI?
Yes, if you want compact, power-efficient hardware. Mini-PCs with Ryzen AI Max or Apple Silicon pack plenty of memory into a small footprint. Make sure to choose enough RAM upfront, as mini-PCs often aren’t upgradeable.
Where do I find buying advice?
The AI Hardware Buying Guide section has recommendations for beginner systems. The Beginner AI PC article offers concrete models and price ranges.
What is quantization and how does it relate to hardware?
Quantization reduces model precision so it runs with less memory. This lets larger models fit on smaller hardware. Read more in the Quantization article.


