Hardware Requirements for AI Agents
What this article covers
- What hardware you need for AI agents.
- GPU, VRAM, RAM, and CPU requirements.
- Recommendations for different setups (entry-level to high-end).
- How to size hardware for your agents.
- Best practices for cost-efficiency and scaling.
Introduction: Understanding hardware requirements
AI agents need computing power: a GPU for the model, RAM for agent code, and a CPU for tools and orchestration. Requirements depend on the model size, number of agents, and complexity of tasks.
This article is for anyone planning hardware for AI agents. For foundational concepts, see Hardware and VRAM Calculator.
Why do I need to understand hardware requirements?
Imagine buying a GPU for your agent. Too small: the agent runs slowly and frustrates users. Too large: expensive and wasteful. The right hardware depends on your model (7B or 70B?), how many agents you’re running (1 or 10?), and your workload (light or heavy?).
Hardware requirements at a glance
GPU plus VRAM for the model is the biggest factor. RAM handles agent code and tools. CPU manages orchestration. For a single agent, an RTX 3060 works. For 10 agents, you’ll want an RTX 4090 or multiple GPUs.
The core principle: GPU is your bottleneck, so size it correctly.
Who should read this?
- Buyers planning hardware purchases for agents.
- Self-hosters setting up servers for agents.
- Decision-makers budgeting for hardware.
- Tinkerers maximizing existing hardware.
Key terms
- GPU - Graphics card for inference. Useful for LLMs.
- VRAM - GPU memory. Useful for model size.
- RAM - System memory. Useful for agent code.
- CPU - Processor. Useful for tools and orchestration.
- VRAM Calculator - Memory requirements. Useful for planning.
- Ollama - Model server. Useful for serving models.
Hardware requirements by model size
| Model | VRAM (Q4) | VRAM (Q8) | RAM | CPU | GPU recommendation |
|---|---|---|---|---|---|
| 3B | 2 GB | 3 GB | 4 GB | 4 Cores | GTX 1650 |
| 7B | 4-5 GB | 7 GB | 8 GB | 4 Cores | RTX 3060 |
| 8B | 5 GB | 8 GB | 8 GB | 4 Cores | RTX 3060 |
| 13B | 8 GB | 13 GB | 16 GB | 6 Cores | RTX 4070 |
| 14B | 9 GB | 14 GB | 16 GB | 6 Cores | RTX 4070 |
| 32B | 20 GB | 32 GB | 32 GB | 8 Cores | RTX 4090 |
| 70B | 40 GB | 70 GB | 64 GB | 16 Cores | 2x RTX 4090 |
Setup recommendations
Entry-level: 1 agent, small model
CPU: 4 Cores (Intel i5 / AMD Ryzen 5)
RAM: 16 GB
GPU: RTX 3060 12GB
VRAM: 12 GB
Model: llama3.1:8b (Q4)
Cost: ~800-1,000 €
Mid-range: 2-3 agents, medium model
CPU: 8 Cores (Intel i7 / AMD Ryzen 7)
RAM: 32 GB
GPU: RTX 4070 12GB
VRAM: 12 GB
Model: llama3.1:8b (Q4) or 13B (Q4)
Cost: ~1,500-2,000 €
High-end: 5-10 agents, large model
CPU: 16 Cores (Intel i9 / AMD Ryzen 9)
RAM: 64 GB
GPU: RTX 4090 24GB
VRAM: 24 GB
Model: llama3.1:70b (Q4) or multiple 8B models
Cost: ~3,000-4,000 €
Enterprise: Many agents, horizontal scaling
CPU: 32 Cores (AMD EPYC / Intel Xeon)
RAM: 128 GB
GPU: 2x RTX 4090 or A100
VRAM: 48 GB
Model: Multiple models or 70B
Cost: ~10,000+ €
CPU vs. GPU for agents
| Component | Function | Importance |
|---|---|---|
| GPU | LLM inference | ⭐⭐⭐⭐⭐ Critical |
| VRAM | Model storage | ⭐⭐⭐⭐⭐ Critical |
| RAM | Agent code, tools | ⭐⭐⭐⭐ Important |
| CPU | Orchestration, tools | ⭐⭐⭐ Medium |
| SSD | Model loading, logs | ⭐⭐ Nice to have |
Practical example: Planning hardware
def plan_hardware(agents, model_size, parallel_requests):
"""Plan hardware requirements"""
# VRAM per agent
vram_per_agent = {
"7b": 5,
"8b": 5,
"13b": 8,
"32b": 20,
"70b": 40
}
# Total VRAM needed
total_vram = agents * vram_per_agent[model_size]
# RAM for agent code
ram_per_agent = 2 # GB
total_ram = agents * ram_per_agent + 8 # +8 for system
# CPU for orchestration
cpu_cores = max(4, agents)
# GPU recommendation
if total_vram <= 12:
gpu = "RTX 3060/4070 (12GB)"
elif total_vram <= 24:
gpu = "RTX 4090 (24GB)"
elif total_vram <= 48:
gpu = "2x RTX 4090 or A100"
else:
gpu = "Multiple servers or cloud"
return {
"gpu": gpu,
"vram_needed": total_vram,
"ram_needed": total_ram,
"cpu_cores": cpu_cores
}
# Example: 5 agents, 8B model
result = plan_hardware(agents=5, model_size="8b", parallel_requests=5)
# Output: GPU: RTX 4090, VRAM: 25 GB, RAM: 18 GB, CPU: 8 Cores
Important warnings
- Hardware failure: GPUs and servers can fail. You need a backup strategy.
- Power costs: Continuous GPU use costs 50-100 €/month in electricity.
- Cooling: GPUs need proper cooling. Overheating reduces lifespan.
- Physical space: Servers need room and adequate ventilation.
Common pitfalls
- GPU too small: Insufficient VRAM means the model won’t load or loads very slowly.
- Insufficient RAM: Agent code and tools need RAM. 16 GB is the bare minimum.
- Underestimating CPU: Multiple agents need CPU for orchestration.
- Forgetting power costs: Continuous GPU operation is expensive.
- No headroom: Plan spare capacity for growth.
Further reading
- Hardware - Hardware overview.
- VRAM Calculator - Calculate memory needs.
- GPU recommendations - GPU buying guide.
- Running AI agents locally - Overview.
- vLLM - For performance.
- Multiple agents - For scaling.
Key takeaways:
- GPU and VRAM are the most important factors for AI agents.
- 8B model: ~5 GB VRAM. 70B: ~40 GB VRAM.
- RAM for agent code: 2 GB per agent plus 8 GB for the system.
- CPU for orchestration: 4-16 cores depending on agent count.
- For 1 agent: RTX 3060. For 10 agents: RTX 4090 or multiple GPUs.
FAQ
What hardware do I need?
How much VRAM do I need?
Which GPU do you recommend?
How important is the CPU?
How much RAM do I need?
How many agents can I run?
What does this hardware cost?
How do I scale hardware?
Sources and further reading
- Ollama - Model server.
- NVIDIA - GPUs.
- VRAM Calculator - Calculate memory requirements.


