AI Hardware
What this article covers
- Why matching hardware matters for local AI.
- What AI hardware content is available on BotServ.
- Links to foundational material, buying guides, and related topics.
Introduction
Running local AI requires hardware that fits your workload. RAM, VRAM, processor, and storage determine which models you can run and how smoothly they perform. This section provides foundational knowledge and practical buying recommendations for newcomers and experienced users alike.
Why do I need the right AI hardware?
A language model is a large file that gets loaded into memory and processed there. Without enough RAM or VRAM, smaller or less quantized models simply won’t fit. With proper hardware, Ollama, LM Studio, or an AI agent runs smoothly instead of waiting endlessly on computations.
Here’s a concrete example: running Llama 3.1 8B requires roughly 8 GB of storage for the quantized model. On a GPU with 8 GB VRAM, it runs smoothly. On a machine with only 16 GB RAM and CPU inference, each response takes several seconds. Knowing this upfront saves both money and frustration.
AI hardware at a glance
Local AI uses two inference paths: CPU-based with the model in RAM, or GPU-based with the model in VRAM. GPU inference is significantly faster but demands a graphics card with sufficient memory. An APU like Apple Silicon or Ryzen AI Max combines CPU and GPU on a single chip and shares memory between them.
Who is this section for?
- You want to run local AI and understand what your machine can handle.
- You’re considering buying a new desktop, mini-PC, or graphics card.
- You’re new to this and want to evaluate marketing claims more critically.
- You don’t need deep prior knowledge. Technical terms are explained inline in the articles.
Content and articles
- Fundamentals - CPU, GPU, RAM, VRAM, SSD, APU, TOPS, and how they work together.
- Buying guide - What to look for when purchasing, with component tables and terminology.
- Starter AI PCs - Detailed buying guidance with price ranges, variants, and specific models.
Key terms
- TOPS/NTOPS - Tera Operations Per Second. The unit of measurement for AI compute performance. Higher values mean a chip can run models faster.
- APU - Accelerated Processing Unit: CPU and GPU on one chip, as in Apple Silicon or Ryzen AI Max.
- Unified/Shared Memory - System RAM is shared by both CPU and GPU. Critical for Apple Silicon and Ryzen AI Max.
- Quantization - Reduces model precision so it runs with less memory.
- VRAM - Memory on a dedicated graphics card. Essential for GPU inference.
Recommended mini-PCs and GPUs
You’ll find suitable mini-PCs, Ryzen AI Max devices, and GPUs for local AI in the Amazon shop. I suggest checking back regularly since prices and availability shift quickly.
Mini-PCs for local AI in the Amazon shop
Bei Amazon ansehenAffiliate-Link: Bei einem Kauf erhalten wir möglicherweise eine Provision.
Common pitfalls
- RAM is not VRAM - 32 GB of RAM help little if your GPU has only 4 GB VRAM. For GPU inference, what matters is graphics memory.
- TOPS don’t tell the whole story - A high TOPS value doesn’t automatically mean your model will run fast. Memory bandwidth and architecture matter equally.
- Mini-PCs aren’t upgradable - If you want more storage later, you often need a new device. Choose enough RAM from the start.
Further reading
FAQ - Frequently asked questions
What’s the difference between RAM and VRAM?
RAM is your computer’s system memory, VRAM is your graphics card’s memory. With CPU inference, the model sits in RAM. With GPU inference, it sits in VRAM. Learn more in the AI hardware fundamentals.
Do I need a graphics card for local AI?
Not necessarily. Small models like Llama 3.1 8B run on CPU. For faster responses and larger models, a GPU with adequate VRAM is recommended.
What does TOPS mean?
TOPS stands for Tera Operations Per Second. It measures how many computational operations a chip can perform per second for AI tasks. Higher values suggest potentially faster AI inference, but memory bandwidth and architecture are equally important.
How much RAM do I need for local AI?
For 7B models, at least 16 GB RAM, preferably 32 GB. Larger models like 70B require 64 GB or more.
How much VRAM do I need for local AI?
For quantized 7B models, at least 8 GB VRAM. For 13B models, 12 GB or more. 70B models need 24 GB or more.
What is an APU?
An APU combines CPU and GPU on a single chip. Examples include Apple Silicon and AMD Ryzen AI Max. The advantage is that both processors can access the same memory pool.
Is a mini-PC worth it for local AI?
Yes, if you want compact, power-efficient hardware. Mini-PCs with Ryzen AI Max or Apple Silicon pack substantial memory into a small form factor. Just make sure to choose enough RAM upfront.
Where can I find buying guidance?
Check the buying guide for starter system recommendations. The Starter AI PCs article offers concrete models and price points.
Why is an SSD important?
Language models can be several gigabytes in size. An NVMe SSD loads the model much faster than an HDD. This reduces startup time and enables smoother operation.
What is Unified Memory?
With Unified Memory, CPU and GPU share the same system RAM. This means a Mac with 32 GB RAM can use all 32 GB for GPU tasks. A traditional discrete graphics card has separate, limited VRAM.


