Skip to content
BotServBotServ
AI HardwareBasicsRAMVRAMGPUCPUSSDTOPS

AI Hardware Basics

AI Hardware Basics: CPU, GPU, RAM, VRAM, SSD, APU and TOPS explained. What you need for local AI and how components work together.

S

schutzgeist

6 min read
AI Hardware Basics

AI Hardware Fundamentals

What This Article Covers

  • Which hardware components matter for local AI and why.
  • How CPU, GPU, RAM, VRAM, and SSD work together.
  • What TOPS, APU, and Unified Memory mean.
  • Links to all fundamentals articles and buying guides.

Introduction

Running AI models locally requires the right hardware. A language model is a large file that gets loaded into memory and processed there. Whether it runs smoothly depends on your CPU, GPU, RAM, VRAM, and storage. This section gathers foundational knowledge to help you understand, assemble, or purchase an AI PC.

Why Learn AI Hardware Basics?

Without this knowledge, you’re buying blind. You risk ending up with a machine that’s either too weak for your models or too expensive for what you actually need. Take this example: running Llama 3.1 8B requires roughly 8 GB of memory for the quantized model. On a GPU with 8 GB VRAM, it runs smoothly. On a machine with only 16 GB RAM using CPU inference, each response takes several seconds. Knowing this upfront saves money and frustration.

These basics also help you evaluate marketing claims. A mini-PC with 128 GB RAM sounds impressive, but if memory bandwidth is low, a large model still runs slowly. TOPS values on a datasheet don’t tell the whole story because they measure peak performance under ideal conditions.

AI Hardware Explained Simply

Local AI loads a model into memory and computes responses from it. There are two approaches:

  • CPU Inference: The model sits in regular system RAM. The CPU calculates the responses. Slower, but flexible and cheap.
  • GPU Inference: The model sits in graphics memory (VRAM). The GPU calculates the responses in parallel, much faster.

Memory capacity determines which models fit at all. Compute power determines how fast you get answers. An SSD ensures the model loads quickly. An APU like Apple Silicon or Ryzen AI Max combines CPU and GPU on one chip and shares the same memory pool.

Who Should Read This?

  • You want to run AI locally and understand what your machine can do.
  • You’re considering buying a new PC or mini-PC.
  • You’re new to this and want to make sense of marketing claims.
  • You need no prior knowledge. Technical terms are explained as they come up.

Key Components

ComponentRoleWhat to Look For
CPURuns inference and operating systemModern cores, plenty of cache
RAMStores models during CPU inference16 GB minimum, 32 GB better, 64-128 GB ideal
GPU/VRAMFast parallel computation8 GB minimum, 16 GB or more preferred
APUCPU and GPU on one chipUnified or Shared Memory uses common RAM
SSDFast storage for models500 GB or more, preferably NVMe

Key Terms

  • TOPS/NTOPS - Tera Operations Per Second. The unit of measurement for AI compute power. Higher values mean faster model execution.
  • APU - Accelerated Processing Unit: CPU and GPU on one chip, like Apple Silicon or Ryzen AI Max.
  • Unified/Shared Memory - RAM is used jointly by CPU and GPU. Essential for Apple Silicon and Ryzen AI Max.
  • Quantization - Reduces model precision to run it with less memory. Lets larger models fit on smaller hardware.
  • VRAM - Memory on a dedicated graphics card. Critical for GPU inference.
  • LLM - Large Language Model.

Content and Articles

Related articles:

Common Pitfalls

  • RAM is not VRAM - 32 GB RAM helps little if your GPU has only 4 GB VRAM. For GPU inference, graphics memory is what counts.
  • TOPS don’t tell the whole story - A high TOPS value doesn’t guarantee your model runs fast. Memory bandwidth and architecture matter equally.
  • Mini-PCs aren’t always upgradeable - If you want more memory later, you often need a new device. Choose enough RAM from the start.
  • Apple Silicon is efficient, but not everything runs natively - Some open-source models are optimized for CUDA first and run better on Linux or Windows with an Nvidia GPU.

Further Reading

FAQ - Common Questions

What’s the difference between RAM and VRAM?

RAM is your machine’s main memory, VRAM is the graphics card’s memory. During CPU inference, the model sits in RAM. During GPU inference, it sits in VRAM. GPU inference is much faster but needs enough VRAM. Read more in RAM vs. VRAM.

What matters more: RAM or VRAM?

Both. CPU-based inference needs lots of RAM, GPU-based inference primarily needs VRAM. For larger models, a graphics card with sufficient VRAM is usually faster.

What does TOPS mean?

TOPS stands for Tera Operations Per Second. It measures how many compute operations a chip can perform per second for AI tasks. Higher values mean potentially faster AI inference, but memory bandwidth and architecture are equally important.

Do I need a graphics card for local AI?

Not necessarily. Smaller models like Llama 3.1 8B run on the CPU alone. For faster responses and larger models, a GPU with sufficient VRAM is recommended.

What is an APU?

An APU combines CPU and GPU on one chip. Examples include Apple Silicon and AMD Ryzen AI Max. The advantage is both processors can access the same memory, which helps significantly with local AI.

What is Unified Memory?

With Unified Memory, CPU and GPU share the same system RAM. This means a Mac with 32 GB RAM can use all 32 GB for the GPU as well. With a traditional graphics card, VRAM is separate and limited.

How much RAM do I need for local AI?

For 7B models, at least 16 GB RAM; 32 GB is better. Larger models like 70B need 64 GB or more. With cloud APIs, you only need memory for the framework itself.

How much VRAM do I need for local AI?

For quantized 7B models, at least 8 GB VRAM. For 13B models, 12 GB or more. 70B models need 24 GB or more, often multiple GPUs.

Why is an SSD important?

A language model can be several gigabytes in size. An NVMe SSD loads the model much faster than an HDD. This reduces startup time and keeps your workflow smooth.

Is a mini-PC worth it for local AI?

Yes, if you want compact, power-efficient hardware. Mini-PCs with Ryzen AI Max or Apple Silicon pack plenty of memory into a small footprint. Make sure to choose enough RAM upfront, as mini-PCs often aren’t upgradeable.

Where do I find buying advice?

The AI Hardware Buying Guide section has recommendations for beginner systems. The Beginner AI PC article offers concrete models and price ranges.

What is quantization and how does it relate to hardware?

Quantization reduces model precision so it runs with less memory. This lets larger models fit on smaller hardware. Read more in the Quantization article.

Back to Blog
Share:

Related Posts