Skip to content
BotServBotServ
OllamaRAGVector DatabaseEmbeddingsChromaDBQdrant

Building RAG with Ollama

Local Retrieval-Augmented Generation with Ollama, embeddings, and vector databases. Chunking, search, and prompts.

S

schutzgeist

4 min read
Building RAG with Ollama

Building RAG with Ollama

What this article covers

  • What RAG is and how it works
  • The components your local RAG system needs
  • How to split documents into chunks
  • How embeddings and vector databases work together
  • A practical Python example using Ollama and ChromaDB

Introduction: Building RAG with Ollama

Retrieval-Augmented Generation, or RAG, extends language models with external knowledge. Instead of relying only on what the model learned during training, relevant document sections are retrieved from a knowledge base at runtime and inserted into the prompt. This approach makes responses more accurate, more current, and enables the model to access your own documents. Ollama is ideal for RAG because both language models and embedding models can run locally.

This article walks you through building a simple RAG system using Ollama, an embedding model, and a vector database.

Key terms

  • RAG: Retrieval-Augmented Generation
  • Embedding: Vector representation of text
  • Vector database: Storage for embeddings with similarity search
  • Chunk: A small text segment extracted from a document
  • Retriever: The component that finds relevant chunks
  • Context: The assembled chunks passed to the language model
  • Prompt: The user query together with context
  • Cosine similarity: A metric for measuring the distance between two vectors

Architecture of a local RAG system

A typical RAG system has four main stages:

  1. Load documents: Read PDFs, text files, Markdown, or web pages
  2. Chunking: Split documents into overlapping pieces
  3. Generate embeddings: Convert each chunk into a vector
  4. Store and search: Place vectors in a database and query them
  5. Generate response: Send relevant chunks as context to Ollama

Loading documents

Use Python with langchain or simple file operations:

from pathlib import Path

doc_path = Path("dokumente/handbuch.md")
text = doc_path.read_text(encoding="utf-8")

For PDFs, consider pymupdf or pdfplumber.

Chunking

Split documents into fixed-size pieces with overlap:

def chunk_text(text, chunk_size=500, overlap=50):
    chunks = []
    for i in range(0, len(text), chunk_size - overlap):
        chunks.append(text[i:i + chunk_size])
    return chunks

chunks = chunk_text(text)

Common parameters:

  • Chunk size: 300 to 1000 characters
  • Overlap: 10 to 20 percent

Semantic chunking, such as splitting by paragraphs or headings, often works better than counting characters alone.

Generating embeddings

import requests
import json

def embed(texts, model="nomic-embed-text"):
    url = "http://localhost:11434/api/embed"
    payload = {"model": model, "input": texts}
    resp = requests.post(url, json=payload)
    return resp.json()["embeddings"]

vectors = embed(chunks)

Setting up a vector database

ChromaDB can run locally:

pip install chromadb
import chromadb

client = chromadb.Client()
collection = client.create_collection(name="wissen")

for i, (chunk, vector) in enumerate(zip(chunks, vectors)):
    collection.add(
        ids=[str(i)],
        embeddings=[vector],
        documents=[chunk]
    )

Searching

question = "Wie wird das Passwort zurückgesetzt?"
question_vector = embed([question])[0]

results = collection.query(
    query_embeddings=[question_vector],
    n_results=3
)

context = "\n\n".join(results["documents"][0])

Generating the response

import requests

url = "http://localhost:11434/api/chat"
payload = {
    "model": "llama3.1",
    "messages": [
        {"role": "system", "content": "Beantworte die Frage ausschliesslich anhand des Kontexts."},
        {"role": "user", "content": f"Kontext:\n{context}\n\nFrage: {question}"}
    ],
    "stream": False
}

resp = requests.post(url, json=payload)
print(resp.json()["message"]["content"])

System prompt for RAG

Your system prompt should instruct the model to answer only from the provided context:

Du beantwortest Fragen ausschliesslich anhand des bereitgestellten Kontexts.
Wenn die Antwort nicht im Kontext steht, ehrlich angeben, dass keine passende Information vorliegt.

Alternative vector databases

DatabaseHighlights
ChromaDBSimple, great for prototyping
QdrantFast, scalable, REST API
WeaviateModel-agnostic, strong for enterprise
pgvectorPostgreSQL extension
FAISSMeta library, extremely fast

Tips for effective RAG

  • Clean and preprocess documents thoroughly
  • Choose sensible chunk sizes
  • Consider semantic chunking
  • Use an embedding model matched to your domain and language
  • Adjust the number of chunks returned
  • Monitor responses for hallucinations
  • Test and iterate on questions and answers

Common pitfalls

  • Chunks too large: Too much irrelevant information in context
  • Chunks too small: Important connections get lost
  • Wrong embedding model: Language or domain mismatch
  • Unclean documents: Headers, footers, or HTML tags interfere
  • Too many chunks in context: Exceeds the model’s context window
  • No source attribution: Users cannot verify answers
  • Hallucinations: Model invents answers despite having context

Extensions

  • Re-ranking: First retrieve many chunks quickly, then sort by relevance
  • Hybrid search: Combine keyword and vector search
  • Metadata: Filter by date, source, or category
  • LangChain / LlamaIndex: Frameworks that simplify RAG pipelines
  • Open WebUI: Offers built-in RAG functionality for Ollama

Further reading and resources

FAQ: RAG with Ollama

Do I need a GPU for RAG? Recommended, but not required. Embeddings and inference are significantly faster on GPU.

Which embedding model should I use? nomic-embed-text is a solid general-purpose choice; mxbai-embed-large handles multiple languages better.

Can I use PDFs? Yes, after text extraction and chunking.

How large should chunks be? 300 to 1000 characters, depending on your documents and questions.

Is Open WebUI suitable for RAG? Yes, it includes built-in RAG functionality with document upload.

Sources and further reading

Summary: Building RAG with Ollama

RAG extends Ollama with your own documents and current knowledge. The architecture covers loading, chunking, embedding, vector storage, and generative responses. With ChromaDB, Python, and Ollama, you can build a prototype RAG system quickly. The keys are clean documents, sensible chunk sizes, a matching embedding model, and a well-written system prompt. By testing RAG iteratively, you improve answer quality precisely and with strong privacy guarantees.

Back to Blog
Share:

Related Posts