Skip to content
BotServBotServ
RAGAI PCBuying GuideVector DatabaseEmbeddingVRAMLocal AIAmazon

PC for RAG: Hardware for Knowledge Databases

Best PC for local RAG? Hardware for vector databases, embedding models, and LLM inference. VRAM, RAM, and SSD recommendations with builds.

S

schutzgeist

11 min read
PC for RAG: Hardware for Knowledge Databases

PC for RAG: Hardware for Knowledge Bases

What this article covers

  • Which hardware components truly matter for local RAG, from GPU to SSD
  • How much VRAM you need when an LLM and embedding model run simultaneously
  • Why vector databases like Chroma and Qdrant consume significant RAM
  • Which builds suit RAG, from budget systems to workstations
  • Common pitfalls to avoid when building a RAG system

Introduction: PC for RAG explained

Local RAG combines a language model with your own knowledge database. Instead of training the model on everything, you search your documents at runtime and feed relevant passages as context to the LLM. It runs locally, no cloud, no API costs.

The catch: RAG places different hardware demands than pure LLM inference. You’re running an LLM, an embedding model, and a vector database simultaneously. All three compete for VRAM, RAM, and memory bandwidth. A PC that runs a 7B model smoothly can struggle with RAG once you add a vector database and embedding model.

This article explains what hardware makes sense for RAG and which builds fit different budgets. For additional recommendations, check out our buying guide.

Why do I need specialized hardware for RAG?

A standard LLM setup loads one model into VRAM and generates text. RAG does more: when you ask a question, the system first searches your vector database for matching text passages. An embedding model converts your question into a vector. The vector database compares this vector against all stored vectors and returns the best matches. Only then does the assembled prompt go to the LLM.

Three things happen simultaneously or in sequence:

  1. Embedding model: Converts the query into vectors, needs VRAM or RAM
  2. Vector database: Holds all embeddings in memory, needs RAM
  3. LLM: Generates the answer from context, needs VRAM

If you run a 7B LLM with 6 GB VRAM and an embedding model like nomic-embed-text needs another 1 to 2 GB, that’s already 8 GB consumed. Add the KV cache, which quickly reaches several GB with longer documents. Skip planning for this and you’ll hit performance walls.

The fundamentals are covered in our article Sizing hardware correctly.

RAG PC at a glance

For RAG you need three things at once: VRAM for the LLM and embedding model, RAM for the vector database, and a fast SSD. As a rough rule: at least 12 GB VRAM for a 7B LLM plus embedding model, 32 GB RAM for smaller vector databases, and an NVMe SSD. If you’re searching larger collections, plan for 64 GB RAM or more.

Who is this article for?

This article targets beginners and experienced users who are planning or upgrading a PC for local RAG. If you understand what a vector database does and what VRAM is, you’re ready. Newcomers should start with RAG basics for a broader overview.

Key RAG hardware terms

TermMeaning
VRAMGPU video memory. Space for the LLM and embedding model.
RAMSystem memory. Holds the vector database in memory.
Embedding modelConverts text to vectors. Needs VRAM or RAM depending on setup.
Vector databaseStores embeddings and enables similarity search. Examples: Chroma, Qdrant.
SSDDrive for models and documents. NVMe is significantly faster than SATA.
NVMeFast SSD interface. Important for short load times at startup.
KV cacheStores computed attention values during inference. Grows with context length.
ChunkingDivision of documents into passages before embedding.
RetrievalSearch operation in the vector database for matching text passages.
RerankingSecond scoring of search results for better relevance. Requires additional compute.

RAG hardware requirements

RAG means three components need resources simultaneously. The table below shows typical requirements for a local RAG setup with a 7B LLM.

ComponentVRAM needRAM needNotes
7B LLM (Q4)5 to 6 GB16 GBMain model for answer generation
Embedding model1 to 2 GB4 GBe.g. nomic-embed-text, all-MiniLM
KV cache (8k context)1 to 3 GB-Grows with context length
Vector database-4 to 32 GBDepends on number of vectors
Operating system + other-8 GBBrowser, IDE, other programs

Combined, you’re looking at roughly 8 to 11 GB VRAM and 32 to 64 GB RAM for a usable 7B RAG setup. Using a 13B LLM requires proportionally more VRAM. Details are in our articles on embedding models and vector databases.

GPU for RAG

The GPU in a RAG system must load two models: the LLM and the embedding model. Both need VRAM. If both don’t fit in VRAM together, you have two strategies.

Strategy 1: Both models in VRAM simultaneously. This is fastest. On each query, embedding and search happen instantly without load time. You need enough VRAM for both plus KV cache. A 7B LLM in Q4 (6 GB) plus nomic-embed-text (1 GB) plus KV cache (2 GB) needs around 9 GB VRAM. An RTX 3060 with 12 GB handles this.

Strategy 2: Load models sequentially. If VRAM is tight, you unload the LLM, load the embedding model, embed the query, and reload the LLM. This costs several seconds of load time each cycle. For interactive use it’s inconvenient, for batch processing it’s acceptable.

With Ollama you can manage both models. Ollama keeps models in memory when space allows and automatically unloads them when needed. The OLLAMA_KEEP_ALIVE setting controls how long models stay loaded.

Recommended GPUs for RAG:

  • RTX 3060 12 GB: Entry point for 7B RAG. Both models fit in VRAM simultaneously.
  • RTX 4070 12 GB: Faster, same VRAM. Good for 7B with longer contexts.
  • RTX 3090 24 GB: Plenty of VRAM for 13B LLM plus embedding model plus large KV cache.
  • RTX 4090 24 GB: Maximum speed for 13B RAG setups.

RAM and vector database

Vector databases like Chroma and Qdrant keep embeddings in RAM to enable fast similarity search. The more documents you have, the more RAM you need.

An embedding with 768 dimensions takes roughly 3 KB of storage. With 100,000 documents split into 10 chunks each, that’s 1 million vectors, or about 3 GB RAM just for vectors. Add metadata and index structures, which roughly doubles the requirement.

DocumentsChunksVectorsRAM needed (approx.)
1,00010,00010,0000.1 GB
10,000100,000100,0000.8 GB
100,0001,000,0001,000,0008 GB
500,0005,000,0005,000,00040 GB

If you’re searching large document collections, you’ll need significantly more RAM. 32 GB is sufficient for small to medium collections. With several hundred thousand documents, plan for 64 GB or more.

SSD: Why NVMe Matters

RAG continuously loads models and documents. When Ollama starts, it loads both the LLM and the embedding model. An NVMe SSD with 3,500 MB/s or faster loads a 5 GB model in under two seconds. A SATA SSD with 550 MB/s takes about nine seconds.

Vector databases benefit from NVMe too. At startup, Chroma or Qdrant loads the index from disk into RAM. With large collections, this can be several gigabytes. NVMe cuts startup time noticeably.

You’ll also store original documents on the SSD. Large collections quickly reach several hundred gigabytes. A 2 TB NVMe SSD gives you enough headroom for models, vector database, and documents.

Hardware für RAG im Amazon Shop

Bei Amazon ansehen

Affiliate-Link: Bei einem Kauf erhalten wir möglicherweise eine Provision.

Build Suggestions for RAG PCs

Build 1: Budget, Around 700-900 EUR

  • CPU: AMD Ryzen 5 5600 or Intel Core i5-12400F
  • RAM: 32 GB DDR4
  • GPU: RTX 3060 12 GB
  • Storage: 1 TB NVMe SSD
  • Power Supply: 550 W, 80+ Bronze

This build runs a 7B LLM plus embedding model simultaneously. 32 GB RAM handles vector databases up to about 100,000 vectors. Ideal for getting started.

Build 2: Mid-Range, Around 1,400-2,000 EUR

  • CPU: AMD Ryzen 5 7600 or Intel Core i5-13600K
  • RAM: 64 GB DDR5
  • GPU: RTX 4070 12 GB or used RTX 3090 24 GB
  • Storage: 2 TB NVMe SSD
  • Power Supply: 750 W, 80+ Gold

With 24 GB VRAM (RTX 3090), you run 13B LLM plus embedding model plus large KV cache. 64 GB RAM gives you room for vector databases with hundreds of thousands of vectors.

Build 3: High-End, Around 3,500-5,500 EUR

  • CPU: AMD Ryzen 9 7950X or Intel Core i9-14900K
  • RAM: 128 GB DDR5
  • GPU: 2x RTX 4090 24 GB or 2x RTX 3090 24 GB
  • Storage: 4 TB NVMe SSD
  • Power Supply: 1,200 W, 80+ Platinum

48 GB VRAM lets you run large LLMs up to 70B in Q4, plus embedding model and reranking model at the same time. 128 GB RAM ensures vector databases with millions of entries run smoothly.

RAG Performance Tips

  • Keep models in VRAM: Leave the LLM and embedding model in VRAM if space allows. This saves load time on every query.
  • Choose quantization wisely: Q4_K_M is a solid default for the LLM. For embedding models, Q8 or FP16 often suffice since they’re small. See Quantisierung for details.
  • Limit context length: The KV cache grows with context size. RAG adds multiple chunks as context. Reduce chunk count or size if VRAM runs tight.
  • Use caching: If you ask the same questions repeatedly, cache embeddings and search results. This saves embedding time and database queries.
  • Batch embeddings: When adding many documents, embed them in batches rather than one by one. This uses the GPU more efficiently.
  • Use reranking sparingly: A reranking model improves search results but needs additional VRAM and compute time. Use it only when retrieval quality isn’t good enough.

Common RAG Hardware Pitfalls

  • Underestimating VRAM: Planning only for the LLM and forgetting the embedding model leaves you short. Always budget for both models plus KV cache.
  • Confusing vector database placement: Chroma and Qdrant run in RAM, not VRAM. The GPU handles the LLM and embedding model.
  • Underestimating RAM: Large vector databases need significant RAM. With 16 GB and 500,000 documents, you hit the limit fast. The system starts swapping.
  • Picking too large an embedding model: Large models like bge-large consume more VRAM than smaller ones like all-MiniLM. Choose based on your hardware.
  • Ignoring LLM context length: RAG adds multiple chunks to the prompt. With 10 chunks at 500 tokens each, you’re at 5,000 tokens of context. The KV cache eats VRAM accordingly.
  • Slow SSD as bottleneck: A SATA SSD noticeably slows loading models and vector database startup. NVMe is practically required for RAG.
  • Undersized power supply: Two RTX 3090s draw over 700 W under load. A weak PSU causes crashes.
  • Small case: Large GPUs don’t fit every case. With dual-GPU setups, measure first.

Hardware, Cost, and Security in RAG

Local RAG means your documents and queries never leave your machine. Confidential documents, internal wikis, and personal notes stay on your disk.

Hardware is a one-time cost; no API charges after that. If you regularly search a knowledge base with thousands of documents, a 1,500 EUR build pays for itself quickly. Factor in power costs: an RTX 4090 draws up to 450 W under load.

For security, download LLMs and embedding models only from trusted sources like Ollama’s official library or Hugging Face with verified uploaders.

Further Reading and RAG Hardware Resources

FAQ: PC for RAG - Common Questions

What GPU do I need for local RAG?

A RTX 3060 with 12 GB VRAM handles a 7B LLM plus embedding model. For 13B models, 24 GB VRAM is recommended, such as an RTX 3090.

Can I run RAG without a GPU?

Yes, RAG works on CPU too. The LLM and embedding model run in RAM. But speed drops noticeably. A GPU is recommended for interactive use.

How much RAM do I need for RAG?

Minimum 32 GB for small RAG setups. With vector databases holding hundreds of thousands of vectors, plan for 64 GB or more. Needs grow with document count.

Is an RTX 3060 enough for RAG?

Yes, the RTX 3060 with 12 GB VRAM is a solid entry point. It runs a 7B LLM and small embedding model together. For 13B models, VRAM gets tight.

Can I run LLM and embedding model on the GPU at the same time?

Yes, if you have enough VRAM. Ollama keeps both models in memory when space is available. With tight VRAM, Ollama unloads models automatically.

Which vector database is best for local RAG?

Chroma is simple to set up and good for beginners. Qdrant performs better and scales for large collections. Both run locally and keep data in RAM.

How much storage space does RAG need?

Models take 3 to 40 GB each. Vector databases need roughly 3 KB per vector plus index overhead. Original documents range from 1 to 10 MB depending on format. A 2 TB NVMe SSD works for most setups.

Do I need an NVMe SSD for RAG?

Recommended, yes. RAG loads models regularly and the vector database at startup. NVMe at 3,500 MB/s cuts load times noticeably versus SATA at 550 MB/s.

What is model swapping in RAG?

Model swapping means unloading the LLM to load the embedding model, then the reverse. This happens automatically when VRAM runs short. The downside is load time at each swap.

How does context length affect VRAM needs in RAG?

In RAG, you add multiple retrieved text chunks to the prompt. Each extra token in context grows the KV cache, which sits in VRAM. With 10 chunks at 500 tokens each, that’s 5,000 tokens of context, which can cost several extra GB of VRAM.

Is reranking worth it for local RAG?

Reranking improves the order of search results and can boost answer quality. But it needs additional VRAM for another model and takes compute time. For small document collections, it’s often unnecessary. For large collections with many hits, it helps.

Sources and further reading

Back to Blog
Share:

Related Posts