Skip to content
BotServBotServ
Knowledge BaseRAGSemantic SearchQ&ADocuments

Build Knowledge Bases from Documents

Create knowledge bases from documents. RAG, semantic search, Q&A systems and practical examples.

S

schutzgeist

5 min read
Build Knowledge Bases from Documents

Knowledge Base from Documents

What This Article Covers

  • Building a knowledge base from your documents.
  • How RAG, embeddings, and semantic search work.
  • Asking questions against your documents.
  • Practical examples for internal wikis, FAQ systems, and support.
  • Best practices for quality, performance, and maintenance.

Introduction: Knowledge Bases Explained

A knowledge base makes your documents searchable and answerable. Instead of digging through folders, you ask: “How do I submit a vacation request?” and the system finds the right document and returns an answer with sources.

This article is for anyone building a document knowledge base. For foundational concepts, see Local RAG and Document Analysis.

Why Do You Need a Knowledge Base?

Imagine you have 500 documents: manuals, contracts, reports. Instead of searching through all of them, you ask the system: “What does the contract say about termination?” and get back the answer with sources. That’s RAG: Retrieval-Augmented Generation.

Knowledge Base in a Nutshell

Documents → Chunks → Embeddings → Vector Database. On query: Embedding → Similar Chunks → Context → LLM → Answer with Sources. All local with Ollama + Qdrant.

The core idea: search first, then answer.

Who Should Read This?

  • Knowledge managers making documents discoverable.
  • Support teams automating FAQs.
  • Organizations surfacing internal knowledge.
  • Developers building Q&A systems.

Key Concepts

  • RAG - Retrieval-Augmented Generation. When useful: the concept.
  • Embeddings - Text to vector. When useful: for search.
  • Vector Database - Storage. When useful: for embeddings.
  • Chunking - Splitting text. When useful: for better search.
  • Ollama - Model server. When useful: for embeddings and LLM.

Architecture

INDEXING:
Documents (PDF/Word/Text)
    │
    ▼
Extract text → Chunking (500-800 characters)
    │
    ▼
Embeddings (Ollama: multilingual-e5)
    │
    ▼
Vector database (Qdrant/ChromaDB)

QUERY:
Question
    │
    ▼
Embedding → Vector search → Top-K chunks
    │
    ▼
Context + Question → LLM (Ollama: llama3.1)
    │
    ▼
Answer + Sources

Setup: Qdrant + Ollama

version: "3.8"

services:
  qdrant:
    image: qdrant/qdrant:latest
    ports:
      - "6333:6333"
    volumes:
      - qdrant_data:/qdrant/storage

  ollama:
    image: ollama/ollama:latest
    ports:
      - "11434:11434"
    volumes:
      - ollama_data:/root/.ollama

volumes:
  qdrant_data:
  ollama_data:
# Pull embedding model
ollama pull multilingual-e5

# Pull LLM
ollama pull llama3.1

Practical Example: Knowledge Base

import chromadb
import requests
import uuid

class KnowledgeBase:
    def __init__(self):
        self.chroma = chromadb.HttpClient(host="chromadb", port=8000)
        self.collection = self.chroma.get_or_create_collection("wissen")

    def get_embedding(self, text):
        response = requests.post("http://ollama:11434/api/embeddings", json={
            "model": "multilingual-e5",
            "prompt": text
        })
        return response.json()["embedding"]

    def add_document(self, text, source, metadata=None):
        """Index document"""
        chunks = self.chunk_text(text, 500)
        for i, chunk in enumerate(chunks):
            embedding = self.get_embedding(chunk)
            self.collection.add(
                ids=[f"{source}_{i}"],
                documents=[chunk],
                embeddings=[embedding],
                metadatas=[{
                    "source": source,
                    "chunk": i,
                    **(metadata or {})
                }]
            )

    def ask(self, question):
        """Query the knowledge base"""
        # Find similar chunks
        query_embedding = self.get_embedding(question)
        results = self.collection.query(
            query_embeddings=[query_embedding],
            n_results=5
        )

        context = "\n\n".join(results["documents"][0])
        sources = [m["source"] for m in results["metadatas"][0]]

        # Generate answer
        response = requests.post("http://ollama:11434/api/chat", json={
            "model": "llama3.1",
            "messages": [
                {"role": "system", "content": "Answer the question based on the context. Provide sources. If the context is insufficient, say so honestly."},
                {"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"}
            ],
            "stream": False
        })

        return {
            "answer": response.json()["message"]["content"],
            "sources": list(set(sources))
        }

# Usage
kb = KnowledgeBase()
kb.add_document("Vacation request: complete form HR-2024...", "urlaub.pdf")
result = kb.ask("How do I request vacation?")
print(result["answer"])
print(f"Sources: {result['sources']}")

Practical Example: Indexing Documents from a Folder

import os
from pathlib import Path

def index_folder(kb, folder):
    """Index all documents in a folder"""
    for filepath in Path(folder).rglob("*.txt"):
        with open(filepath) as f:
            text = f.read()
        kb.add_document(text, filepath.name)
        print(f"Indexed: {filepath.name}")

# Index all .txt files
index_folder(kb, "./dokumente/")

Improving Quality

Reranking

def rerank_chunks(question, chunks, top_k=3):
    """Rerank chunks by relevance"""
    scores = []
    for chunk in chunks:
        response = requests.post("http://ollama:11434/api/chat", json={
            "model": "llama3.1",
            "messages": [
                {"role": "user", "content": f"Rate the relevance of this text to the question '{question}' on a scale 0-10. Only the number.\n\nText: {chunk}"}
            ],
            "stream": False
        })
        score = float(response.json()["message"]["content"].strip())
        scores.append((score, chunk))

    scores.sort(reverse=True)
    return [c for s, c in scores[:top_k]]

Source Attribution

def format_answer(result):
    """Format answer with sources"""
    answer = result["answer"]
    sources = result["sources"]

    formatted = f"{answer}\n\n---\nSources:\n"
    for s in sources:
        formatted += f"- {s}\n"
    return formatted

Security Notes

  • Access control: Who can query which documents? Filter by permissions.
  • Prompt injection: Documents may contain injections. See Prompt Injection.
  • Data quality: Garbage documents produce garbage answers. See Preparing Documents.
  • Privacy: All data stays local. See Privacy.

Common Pitfalls

  • Chunks too large: Above 1000 characters lose precision. 300-800 is optimal.
  • Too few chunks: Top-5 isn’t always enough. Retrieve Top-20, then rerank to Top-5.
  • Wrong embedding model: For German, use multilingual-e5 or bge-m3.
  • Missing sources: Without sources, answers can’t be verified.
  • Hallucinations: Poor retrieval leads the model to invent answers.

Further Reading

Key Takeaways:

  • Knowledge base workflow: Documents → Embeddings → Vector database → Q&A.
  • Use multilingual-e5 for German documents.
  • Chunking strategy and reranking determine quality.
  • Always include sources, never trust blindly.
  • Stay local with Ollama + Qdrant to keep all data private.

FAQ

What is a knowledge base?

A system that makes documents searchable and answerable. You ask a question, the system finds relevant documents and responds with citations.

How do I build a knowledge base?

Documents → Extract text → Chunking → Embeddings → Vector database. For queries: Embed question → Search → Add context → LLM → Answer.

Which embedding model for German?

multilingual-e5 or bge-m3. Both support German very well. nomic-embed-text is optimized for English only.

How large should chunks be?

300-800 characters with 50-100 character overlap. Chunks that are too small lose context, while those that are too large lose precision.

How do I prevent hallucinations?

Use good retrieval (correct chunks), clear system prompts (“only from context”), require citations, and allow the model to say “I don’t know” when uncertain.

How do I keep the database current?

Set up a workflow to auto-index new documents. For updates, delete old chunks and reindex. See updating your knowledge base.

What does this cost?

Only hardware costs. Ollama, Qdrant, and ChromaDB are open source. No per-document or per-query API fees.

Are my documents safe?

Yes, if you use local tools. All data stays on your server. With cloud APIs, documents leave your machine, so keep confidential documents local.

References and Further Reading

Back to Blog
Share:

Related Posts