Buying a GPU for Local AI
What this article covers
- The role VRAM plays in GPU selection
- How GPU generations and architectures differ
- Which GPUs suit different model sizes
- Considerations for used and consumer GPUs
- Common purchasing criteria and mistakes
Introduction: Buying a GPU for Local AI
A GPU is the most critical component if you want local AI models to run quickly. It accelerates loading and inference for language, image, and embedding models. Choosing the right GPU depends mainly on model size, desired speed, and budget. VRAM is usually the deciding factor.
This article helps you find the right GPU for local AI projects without overspending on inappropriate hardware.
Key concepts
- VRAM: Video memory on the GPU that determines which model sizes fit.
- CUDA: Nvidia’s programming platform, supported by most AI tools.
- ROCm: AMD’s open-source alternative to CUDA.
- Tensor Cores: Specialized compute units for AI on Nvidia RTX and datacenter GPUs.
- Memory bandwidth: Determines how fast data loads into the GPU.
- FP16: Half-precision computation that saves VRAM and is usually sufficient for AI.
- Quantization: Reducing model precision to fit larger models into less VRAM.
Why is VRAM so important?
A model must fit entirely in VRAM. With 4-bit quantization, a rough rule of thumb applies:
- 7B model: roughly 4 to 6 GB VRAM
- 13B model: roughly 8 to 10 GB VRAM
- 30B model: roughly 20 GB VRAM
- 70B model: roughly 40 GB VRAM
If you want to run a 13B model in 4-bit, you need at least 8 GB VRAM. For 30B, 12 GB won’t be enough.
Consumer GPUs for local AI
Nvidia GeForce RTX 3060 12 GB
- Affordable entry point
- 12 GB VRAM supports 7B and smaller 13B models
- Slower than newer generations
Nvidia GeForce RTX 4060 Ti 16 GB
- Modern and power-efficient
- 16 GB VRAM fits 13B and smaller 30B models
- Lower bandwidth than 4070 Ti or 3090
Nvidia GeForce RTX 4070 Ti Super 16 GB
- Strong performance at moderate prices
- 16 GB VRAM with high bandwidth
- Ideal for 13B and 30B models with quantization
Nvidia GeForce RTX 3090 / 3090 Ti 24 GB
- Plenty of VRAM for the price
- Popular among AI enthusiasts
- High power consumption, often available used
Nvidia GeForce RTX 4090 24 GB
- Currently the fastest consumer GPU
- 24 GB VRAM handles 30B and larger models
- Expensive, but extremely fast
AMD Radeon RX 7900 XTX 24 GB
- Ample VRAM at a good price
- ROCm support is improving, but not as mature as CUDA
- Experimental on many tools
Used GPUs
- RTX 2080 Ti with 11 GB: Inexpensive, but lacks newer Tensor Cores
- RTX 3090 with 24 GB: Very popular used, watch for mining history
- Titan X Pascal with 12 GB: Dated, still usable for 7B models
- Quadro P6000 with 24 GB: Professional card, often cheap used, but slow
With used equipment, inspect for dust, fan condition, and potential mining wear.
Server and datacenter GPUs
- Nvidia A100 40/80 GB: Perfect for large models, very expensive
- Nvidia A40 48 GB: Good alternative to A100
- Nvidia L40S 48 GB: Modern inference card
- Nvidia RTX A6000 48 GB: Professional variant with plenty of VRAM
- AMD Instinct MI100/MI210: Usable with ROCm, but smaller ecosystem
Key purchasing criteria
- VRAM size: Primary bottleneck
- Memory bandwidth: Affects inference speed
- CUDA core or stream processor count: More is faster
- Power consumption: Consider PSU and cooling
- Driver support: Nvidia usually best, AMD varies
- Case space: Long or thick GPUs don’t fit everywhere
- PCIe lanes: x16 ideal, x8 often works too
- Price per GB VRAM: Helps with comparison
What to watch with Nvidia
- Install drivers and CUDA Toolkit
nvidia-smishows utilization and VRAM- For containers enable
--gpus all - Ollama, llama.cpp, and vLLM support Nvidia very well
What to watch with AMD
- ROCm must be compatible, not all GPUs are supported
- Older GCN GPUs like HD 7870 XT are not supported by ROCm
- llama.cpp with Vulkan can serve as a fallback but is slower
- Usually better support on Linux than Windows
Apple Silicon
- No dedicated GPU needed, Unified Memory works like VRAM
- 16 GB RAM sufficient for small 7B models
- 32 to 64 GB Unified Memory enable larger models
- Metal backend in llama.cpp and Ollama is well optimized
Price-to-performance recommendations
| Budget | Recommended GPU |
|---|---|
| Up to €300 | RTX 3060 12 GB used |
| €300-600 | RTX 4060 Ti 16 GB |
| €600-1000 | RTX 4070 Ti Super 16 GB |
| €1000-1500 | RTX 3090 24 GB used or RX 7900 XTX |
| Over €1500 | RTX 4090 24 GB or A6000 |
Common purchasing mistakes
- Too little VRAM: 8 GB barely handles 13B models
- Old GPU without Tensor Cores: Missing AI compute power
- AMD without ROCm verification: May not work at all
- Server GPU without cooling: Passively cooled cards need case fans
- High VRAM but narrow bandwidth: Large model fits but runs slow
- Buying a mining GPU: High wear and degradation
Further links and information
- BotServ.de GPU Comparison
- BotServ.de Used Server Hardware
- BotServ.de Proxmox on Mac Pro Trashcan
- BotServ.de Local AI Hardware Requirements
FAQ: GPU for local AI
Do I need a GPU for Ollama? No, Ollama runs on CPU too. But it’s much faster with a GPU.
Which GPU for 7B models? At least 6 GB VRAM, better 8 to 12 GB.
Which GPU for 70B models? At least 40 GB VRAM. An RTX 3090 barely works at 4-bit quantization, A100 or A6000 are ideal.
Is a used GPU worth it? Yes, especially the RTX 3090 offers lots of VRAM for relatively little money.
Are AMD GPUs suitable for AI? With sufficient VRAM and compatible ROCm support, yes, but the ecosystem isn’t as mature as Nvidia’s.
Sources and further reading
- Nvidia CUDA GPUs: https://developer.nvidia.com/cuda-gpus
- AMD ROCm: https://rocm.docs.amd.com/
- Ollama Hardware: https://ollama.com/library
Summary: Buying a GPU for local AI
The right GPU for local AI is selected primarily by VRAM size. Nvidia consumer cards offer the best ecosystem, but AMD and Apple Silicon are also viable options depending on budget and use case. Used cards with plenty of VRAM like the RTX 3090 are popular but should be checked for wear. By keeping VRAM, bandwidth, power consumption, and driver support in mind, you’ll find the right GPU for Ollama, llama.cpp, and other local AI tools.


