Building RAG with Ollama
What this article covers
- What RAG is and how it works
- The components your local RAG system needs
- How to split documents into chunks
- How embeddings and vector databases work together
- A practical Python example using Ollama and ChromaDB
Introduction: Building RAG with Ollama
Retrieval-Augmented Generation, or RAG, extends language models with external knowledge. Instead of relying only on what the model learned during training, relevant document sections are retrieved from a knowledge base at runtime and inserted into the prompt. This approach makes responses more accurate, more current, and enables the model to access your own documents. Ollama is ideal for RAG because both language models and embedding models can run locally.
This article walks you through building a simple RAG system using Ollama, an embedding model, and a vector database.
Key terms
- RAG: Retrieval-Augmented Generation
- Embedding: Vector representation of text
- Vector database: Storage for embeddings with similarity search
- Chunk: A small text segment extracted from a document
- Retriever: The component that finds relevant chunks
- Context: The assembled chunks passed to the language model
- Prompt: The user query together with context
- Cosine similarity: A metric for measuring the distance between two vectors
Architecture of a local RAG system
A typical RAG system has four main stages:
- Load documents: Read PDFs, text files, Markdown, or web pages
- Chunking: Split documents into overlapping pieces
- Generate embeddings: Convert each chunk into a vector
- Store and search: Place vectors in a database and query them
- Generate response: Send relevant chunks as context to Ollama
Loading documents
Use Python with langchain or simple file operations:
from pathlib import Path
doc_path = Path("dokumente/handbuch.md")
text = doc_path.read_text(encoding="utf-8")
For PDFs, consider pymupdf or pdfplumber.
Chunking
Split documents into fixed-size pieces with overlap:
def chunk_text(text, chunk_size=500, overlap=50):
chunks = []
for i in range(0, len(text), chunk_size - overlap):
chunks.append(text[i:i + chunk_size])
return chunks
chunks = chunk_text(text)
Common parameters:
- Chunk size: 300 to 1000 characters
- Overlap: 10 to 20 percent
Semantic chunking, such as splitting by paragraphs or headings, often works better than counting characters alone.
Generating embeddings
import requests
import json
def embed(texts, model="nomic-embed-text"):
url = "http://localhost:11434/api/embed"
payload = {"model": model, "input": texts}
resp = requests.post(url, json=payload)
return resp.json()["embeddings"]
vectors = embed(chunks)
Setting up a vector database
ChromaDB can run locally:
pip install chromadb
import chromadb
client = chromadb.Client()
collection = client.create_collection(name="wissen")
for i, (chunk, vector) in enumerate(zip(chunks, vectors)):
collection.add(
ids=[str(i)],
embeddings=[vector],
documents=[chunk]
)
Searching
question = "Wie wird das Passwort zurückgesetzt?"
question_vector = embed([question])[0]
results = collection.query(
query_embeddings=[question_vector],
n_results=3
)
context = "\n\n".join(results["documents"][0])
Generating the response
import requests
url = "http://localhost:11434/api/chat"
payload = {
"model": "llama3.1",
"messages": [
{"role": "system", "content": "Beantworte die Frage ausschliesslich anhand des Kontexts."},
{"role": "user", "content": f"Kontext:\n{context}\n\nFrage: {question}"}
],
"stream": False
}
resp = requests.post(url, json=payload)
print(resp.json()["message"]["content"])
System prompt for RAG
Your system prompt should instruct the model to answer only from the provided context:
Du beantwortest Fragen ausschliesslich anhand des bereitgestellten Kontexts.
Wenn die Antwort nicht im Kontext steht, ehrlich angeben, dass keine passende Information vorliegt.
Alternative vector databases
| Database | Highlights |
|---|---|
| ChromaDB | Simple, great for prototyping |
| Qdrant | Fast, scalable, REST API |
| Weaviate | Model-agnostic, strong for enterprise |
| pgvector | PostgreSQL extension |
| FAISS | Meta library, extremely fast |
Tips for effective RAG
- Clean and preprocess documents thoroughly
- Choose sensible chunk sizes
- Consider semantic chunking
- Use an embedding model matched to your domain and language
- Adjust the number of chunks returned
- Monitor responses for hallucinations
- Test and iterate on questions and answers
Common pitfalls
- Chunks too large: Too much irrelevant information in context
- Chunks too small: Important connections get lost
- Wrong embedding model: Language or domain mismatch
- Unclean documents: Headers, footers, or HTML tags interfere
- Too many chunks in context: Exceeds the model’s context window
- No source attribution: Users cannot verify answers
- Hallucinations: Model invents answers despite having context
Extensions
- Re-ranking: First retrieve many chunks quickly, then sort by relevance
- Hybrid search: Combine keyword and vector search
- Metadata: Filter by date, source, or category
- LangChain / LlamaIndex: Frameworks that simplify RAG pipelines
- Open WebUI: Offers built-in RAG functionality for Ollama
Further reading and resources
- BotServ.de Ollama Embeddings
- BotServ.de Ollama REST API
- BotServ.de Ollama Commands
- BotServ.de RAG Knowledge Base
- BotServ.de Vector Databases
FAQ: RAG with Ollama
Do I need a GPU for RAG? Recommended, but not required. Embeddings and inference are significantly faster on GPU.
Which embedding model should I use?
nomic-embed-text is a solid general-purpose choice; mxbai-embed-large handles multiple languages better.
Can I use PDFs? Yes, after text extraction and chunking.
How large should chunks be? 300 to 1000 characters, depending on your documents and questions.
Is Open WebUI suitable for RAG? Yes, it includes built-in RAG functionality with document upload.
Sources and further reading
- LangChain RAG: https://python.langchain.com/docs/use_cases/question_answering/
- LlamaIndex: https://www.llamaindex.ai/
- ChromaDB: https://www.trychroma.com/
- Ollama RAG Tutorial: https://github.com/ollama/ollama/blob/main/docs/
Summary: Building RAG with Ollama
RAG extends Ollama with your own documents and current knowledge. The architecture covers loading, chunking, embedding, vector storage, and generative responses. With ChromaDB, Python, and Ollama, you can build a prototype RAG system quickly. The keys are clean documents, sensible chunk sizes, a matching embedding model, and a well-written system prompt. By testing RAG iteratively, you improve answer quality precisely and with strong privacy guarantees.


