PC for RAG: Hardware for Knowledge Bases
What this article covers
- Which hardware components truly matter for local RAG, from GPU to SSD
- How much VRAM you need when an LLM and embedding model run simultaneously
- Why vector databases like Chroma and Qdrant consume significant RAM
- Which builds suit RAG, from budget systems to workstations
- Common pitfalls to avoid when building a RAG system
Introduction: PC for RAG explained
Local RAG combines a language model with your own knowledge database. Instead of training the model on everything, you search your documents at runtime and feed relevant passages as context to the LLM. It runs locally, no cloud, no API costs.
The catch: RAG places different hardware demands than pure LLM inference. You’re running an LLM, an embedding model, and a vector database simultaneously. All three compete for VRAM, RAM, and memory bandwidth. A PC that runs a 7B model smoothly can struggle with RAG once you add a vector database and embedding model.
This article explains what hardware makes sense for RAG and which builds fit different budgets. For additional recommendations, check out our buying guide.
Why do I need specialized hardware for RAG?
A standard LLM setup loads one model into VRAM and generates text. RAG does more: when you ask a question, the system first searches your vector database for matching text passages. An embedding model converts your question into a vector. The vector database compares this vector against all stored vectors and returns the best matches. Only then does the assembled prompt go to the LLM.
Three things happen simultaneously or in sequence:
- Embedding model: Converts the query into vectors, needs VRAM or RAM
- Vector database: Holds all embeddings in memory, needs RAM
- LLM: Generates the answer from context, needs VRAM
If you run a 7B LLM with 6 GB VRAM and an embedding model like nomic-embed-text needs another 1 to 2 GB, that’s already 8 GB consumed. Add the KV cache, which quickly reaches several GB with longer documents. Skip planning for this and you’ll hit performance walls.
The fundamentals are covered in our article Sizing hardware correctly.
RAG PC at a glance
For RAG you need three things at once: VRAM for the LLM and embedding model, RAM for the vector database, and a fast SSD. As a rough rule: at least 12 GB VRAM for a 7B LLM plus embedding model, 32 GB RAM for smaller vector databases, and an NVMe SSD. If you’re searching larger collections, plan for 64 GB RAM or more.
Who is this article for?
This article targets beginners and experienced users who are planning or upgrading a PC for local RAG. If you understand what a vector database does and what VRAM is, you’re ready. Newcomers should start with RAG basics for a broader overview.
Key RAG hardware terms
| Term | Meaning |
|---|---|
| VRAM | GPU video memory. Space for the LLM and embedding model. |
| RAM | System memory. Holds the vector database in memory. |
| Embedding model | Converts text to vectors. Needs VRAM or RAM depending on setup. |
| Vector database | Stores embeddings and enables similarity search. Examples: Chroma, Qdrant. |
| SSD | Drive for models and documents. NVMe is significantly faster than SATA. |
| NVMe | Fast SSD interface. Important for short load times at startup. |
| KV cache | Stores computed attention values during inference. Grows with context length. |
| Chunking | Division of documents into passages before embedding. |
| Retrieval | Search operation in the vector database for matching text passages. |
| Reranking | Second scoring of search results for better relevance. Requires additional compute. |
RAG hardware requirements
RAG means three components need resources simultaneously. The table below shows typical requirements for a local RAG setup with a 7B LLM.
| Component | VRAM need | RAM need | Notes |
|---|---|---|---|
| 7B LLM (Q4) | 5 to 6 GB | 16 GB | Main model for answer generation |
| Embedding model | 1 to 2 GB | 4 GB | e.g. nomic-embed-text, all-MiniLM |
| KV cache (8k context) | 1 to 3 GB | - | Grows with context length |
| Vector database | - | 4 to 32 GB | Depends on number of vectors |
| Operating system + other | - | 8 GB | Browser, IDE, other programs |
Combined, you’re looking at roughly 8 to 11 GB VRAM and 32 to 64 GB RAM for a usable 7B RAG setup. Using a 13B LLM requires proportionally more VRAM. Details are in our articles on embedding models and vector databases.
GPU for RAG
The GPU in a RAG system must load two models: the LLM and the embedding model. Both need VRAM. If both don’t fit in VRAM together, you have two strategies.
Strategy 1: Both models in VRAM simultaneously. This is fastest. On each query, embedding and search happen instantly without load time. You need enough VRAM for both plus KV cache. A 7B LLM in Q4 (6 GB) plus nomic-embed-text (1 GB) plus KV cache (2 GB) needs around 9 GB VRAM. An RTX 3060 with 12 GB handles this.
Strategy 2: Load models sequentially. If VRAM is tight, you unload the LLM, load the embedding model, embed the query, and reload the LLM. This costs several seconds of load time each cycle. For interactive use it’s inconvenient, for batch processing it’s acceptable.
With Ollama you can manage both models. Ollama keeps models in memory when space allows and automatically unloads them when needed. The OLLAMA_KEEP_ALIVE setting controls how long models stay loaded.
Recommended GPUs for RAG:
- RTX 3060 12 GB: Entry point for 7B RAG. Both models fit in VRAM simultaneously.
- RTX 4070 12 GB: Faster, same VRAM. Good for 7B with longer contexts.
- RTX 3090 24 GB: Plenty of VRAM for 13B LLM plus embedding model plus large KV cache.
- RTX 4090 24 GB: Maximum speed for 13B RAG setups.
RAM and vector database
Vector databases like Chroma and Qdrant keep embeddings in RAM to enable fast similarity search. The more documents you have, the more RAM you need.
An embedding with 768 dimensions takes roughly 3 KB of storage. With 100,000 documents split into 10 chunks each, that’s 1 million vectors, or about 3 GB RAM just for vectors. Add metadata and index structures, which roughly doubles the requirement.
| Documents | Chunks | Vectors | RAM needed (approx.) |
|---|---|---|---|
| 1,000 | 10,000 | 10,000 | 0.1 GB |
| 10,000 | 100,000 | 100,000 | 0.8 GB |
| 100,000 | 1,000,000 | 1,000,000 | 8 GB |
| 500,000 | 5,000,000 | 5,000,000 | 40 GB |
If you’re searching large document collections, you’ll need significantly more RAM. 32 GB is sufficient for small to medium collections. With several hundred thousand documents, plan for 64 GB or more.
SSD: Why NVMe Matters
RAG continuously loads models and documents. When Ollama starts, it loads both the LLM and the embedding model. An NVMe SSD with 3,500 MB/s or faster loads a 5 GB model in under two seconds. A SATA SSD with 550 MB/s takes about nine seconds.
Vector databases benefit from NVMe too. At startup, Chroma or Qdrant loads the index from disk into RAM. With large collections, this can be several gigabytes. NVMe cuts startup time noticeably.
You’ll also store original documents on the SSD. Large collections quickly reach several hundred gigabytes. A 2 TB NVMe SSD gives you enough headroom for models, vector database, and documents.
Recommended Hardware on Amazon
Hardware für RAG im Amazon Shop
Bei Amazon ansehenAffiliate-Link: Bei einem Kauf erhalten wir möglicherweise eine Provision.
Build Suggestions for RAG PCs
Build 1: Budget, Around 700-900 EUR
- CPU: AMD Ryzen 5 5600 or Intel Core i5-12400F
- RAM: 32 GB DDR4
- GPU: RTX 3060 12 GB
- Storage: 1 TB NVMe SSD
- Power Supply: 550 W, 80+ Bronze
This build runs a 7B LLM plus embedding model simultaneously. 32 GB RAM handles vector databases up to about 100,000 vectors. Ideal for getting started.
Build 2: Mid-Range, Around 1,400-2,000 EUR
- CPU: AMD Ryzen 5 7600 or Intel Core i5-13600K
- RAM: 64 GB DDR5
- GPU: RTX 4070 12 GB or used RTX 3090 24 GB
- Storage: 2 TB NVMe SSD
- Power Supply: 750 W, 80+ Gold
With 24 GB VRAM (RTX 3090), you run 13B LLM plus embedding model plus large KV cache. 64 GB RAM gives you room for vector databases with hundreds of thousands of vectors.
Build 3: High-End, Around 3,500-5,500 EUR
- CPU: AMD Ryzen 9 7950X or Intel Core i9-14900K
- RAM: 128 GB DDR5
- GPU: 2x RTX 4090 24 GB or 2x RTX 3090 24 GB
- Storage: 4 TB NVMe SSD
- Power Supply: 1,200 W, 80+ Platinum
48 GB VRAM lets you run large LLMs up to 70B in Q4, plus embedding model and reranking model at the same time. 128 GB RAM ensures vector databases with millions of entries run smoothly.
RAG Performance Tips
- Keep models in VRAM: Leave the LLM and embedding model in VRAM if space allows. This saves load time on every query.
- Choose quantization wisely: Q4_K_M is a solid default for the LLM. For embedding models, Q8 or FP16 often suffice since they’re small. See Quantisierung for details.
- Limit context length: The KV cache grows with context size. RAG adds multiple chunks as context. Reduce chunk count or size if VRAM runs tight.
- Use caching: If you ask the same questions repeatedly, cache embeddings and search results. This saves embedding time and database queries.
- Batch embeddings: When adding many documents, embed them in batches rather than one by one. This uses the GPU more efficiently.
- Use reranking sparingly: A reranking model improves search results but needs additional VRAM and compute time. Use it only when retrieval quality isn’t good enough.
Common RAG Hardware Pitfalls
- Underestimating VRAM: Planning only for the LLM and forgetting the embedding model leaves you short. Always budget for both models plus KV cache.
- Confusing vector database placement: Chroma and Qdrant run in RAM, not VRAM. The GPU handles the LLM and embedding model.
- Underestimating RAM: Large vector databases need significant RAM. With 16 GB and 500,000 documents, you hit the limit fast. The system starts swapping.
- Picking too large an embedding model: Large models like
bge-largeconsume more VRAM than smaller ones likeall-MiniLM. Choose based on your hardware. - Ignoring LLM context length: RAG adds multiple chunks to the prompt. With 10 chunks at 500 tokens each, you’re at 5,000 tokens of context. The KV cache eats VRAM accordingly.
- Slow SSD as bottleneck: A SATA SSD noticeably slows loading models and vector database startup. NVMe is practically required for RAG.
- Undersized power supply: Two RTX 3090s draw over 700 W under load. A weak PSU causes crashes.
- Small case: Large GPUs don’t fit every case. With dual-GPU setups, measure first.
Hardware, Cost, and Security in RAG
Local RAG means your documents and queries never leave your machine. Confidential documents, internal wikis, and personal notes stay on your disk.
Hardware is a one-time cost; no API charges after that. If you regularly search a knowledge base with thousands of documents, a 1,500 EUR build pays for itself quickly. Factor in power costs: an RTX 4090 draws up to 450 W under load.
For security, download LLMs and embedding models only from trusted sources like Ollama’s official library or Hugging Face with verified uploaders.
Further Reading and RAG Hardware Resources
- Local RAG - Basics and setup
- RAG Fundamentals - How RAG works
- Embedding Models - Vectors from text
- Vector Databases - Storage and search
- Hardware Buying Guide - More hardware recommendations
- Sizing Hardware Correctly - Planning resources
- Ollama - Manage models locally
- Quantisierung - Make models smaller
FAQ: PC for RAG - Common Questions
What GPU do I need for local RAG?
A RTX 3060 with 12 GB VRAM handles a 7B LLM plus embedding model. For 13B models, 24 GB VRAM is recommended, such as an RTX 3090.
Can I run RAG without a GPU?
Yes, RAG works on CPU too. The LLM and embedding model run in RAM. But speed drops noticeably. A GPU is recommended for interactive use.
How much RAM do I need for RAG?
Minimum 32 GB for small RAG setups. With vector databases holding hundreds of thousands of vectors, plan for 64 GB or more. Needs grow with document count.
Is an RTX 3060 enough for RAG?
Yes, the RTX 3060 with 12 GB VRAM is a solid entry point. It runs a 7B LLM and small embedding model together. For 13B models, VRAM gets tight.
Can I run LLM and embedding model on the GPU at the same time?
Yes, if you have enough VRAM. Ollama keeps both models in memory when space is available. With tight VRAM, Ollama unloads models automatically.
Which vector database is best for local RAG?
Chroma is simple to set up and good for beginners. Qdrant performs better and scales for large collections. Both run locally and keep data in RAM.
How much storage space does RAG need?
Models take 3 to 40 GB each. Vector databases need roughly 3 KB per vector plus index overhead. Original documents range from 1 to 10 MB depending on format. A 2 TB NVMe SSD works for most setups.
Do I need an NVMe SSD for RAG?
Recommended, yes. RAG loads models regularly and the vector database at startup. NVMe at 3,500 MB/s cuts load times noticeably versus SATA at 550 MB/s.
What is model swapping in RAG?
Model swapping means unloading the LLM to load the embedding model, then the reverse. This happens automatically when VRAM runs short. The downside is load time at each swap.
How does context length affect VRAM needs in RAG?
In RAG, you add multiple retrieved text chunks to the prompt. Each extra token in context grows the KV cache, which sits in VRAM. With 10 chunks at 500 tokens each, that’s 5,000 tokens of context, which can cost several extra GB of VRAM.
Is reranking worth it for local RAG?
Reranking improves the order of search results and can boost answer quality. But it needs additional VRAM for another model and takes compute time. For small document collections, it’s often unnecessary. For large collections with many hits, it helps.
Sources and further reading
- Ollama Model Library
- Ollama Documentation
- Chroma Vector Database
- Qdrant Vector Database
- NVIDIA GPU Specifications
- Hugging Face Embedding Models
- Hardware for RAG on Amazon Shop


