Skip to content
BotServBotServ
Document BotRAGPaperless-ngxOllamaCompany ChatbotKnowledge ManagementSelf-HostingPractical Project

Build a Company Document Bot: Paperless-ngx + Ollama + RAG

Create an internal bot answering employee questions from company documents. Paperless-ngx scans, Ollama responds, RAG provides context. Fully local.

S

schutzgeist

4 min read
Build a Company Document Bot: Paperless-ngx + Ollama + RAG

Project: Building a Company Document Bot: Paperless-ngx + Ollama + RAG in Your Own Chat

What This Project Does

Employees ask in chat: “What’s the notice period in the contract with Customer X?” or “Where’s the manual for the machine?” The bot searches the company archive and responds with sources. Completely local: no documents leave the office.

The Stack:

Paper/Email/PDF
    └── Paperless-ngx (Scan + OCR + Archive)
            └── RAG Pipeline (Embeddings + Qdrant)
                    └── Ollama (Answer Model)
                            └── Chat Frontend (Mattermost / Nextcloud Talk / Web)

Prerequisites: A mini PC or server (32 GB RAM is enough), the tools from the Software Stack article, about 2-3 hours for setup.

Why This Architecture

  • Paperless-ngx is the archive backbone: OCR makes scans searchable, tags and correspondents organize content. See Paperless-ngx.
  • RAG instead of training: Documents aren’t baked into the model; they’re searched via embeddings. New documents appear instantly, no retraining needed.
  • Chat as the frontend: Employees ask where they already are: Mattermost or Nextcloud Talk. No new tool to learn.
  • GDPR-compliant: Everything on-premise. Client and contract data stays internal.

Step 1: Paperless-ngx as Archive

# docker-compose.yml (core)
services:
  paperless:
    image: ghcr.io/paperless-ngx/paperless-ngx:latest
    ports: ["127.0.0.1:8000:8000"]
    environment:
      - PAPERLESS_OCR_LANGUAGE=deu
      - PAPERLESS_TIME_ZONE=Europe/Berlin
    volumes:
      - paperless_data:/usr/src/paperless/data
      - ./consume:/usr/src/paperless/consume   # Drop files here → OCR + Archive

volumes:
  paperless_data:

Drop documents into ./consume (or point your scanner directly there): Paperless handles OCR and archival. See Paperless-ngx Integration.

# ingest.py, Index Paperless documents into Qdrant
import requests, uuid
from qdrant_client import QdrantClient
from qdrant_client.models import PointStruct

qdrant = QdrantClient("localhost", port=6333)

def embed(text):
    r = requests.post("http://localhost:11434/api/embeddings",
                      json={"model": "nomic-embed-text", "prompt": text})
    return r.json()["embedding"]

# Paperless API: fetch documents
docs = requests.get("http://localhost:8000/api/documents/",
                    headers={"Authorization": "Token YOUR_TOKEN"}).json()["results"]

for doc in docs:
    text = doc.get("content", "")[:8000]          # OCR text
    vec = embed(text)
    qdrant.upsert("firma", [PointStruct(
        id=str(uuid.uuid4()), vector=vec,
        payload={"title": doc["title"], "id": doc["id"], "text": text[:2000]}
    )])

Use nomic-embed-text as the embedding model (ollama pull nomic-embed-text), Qdrant as the vector database. See Local RAG.

Step 3: The Bot: Question → Retrieval → Answer

# bot.py, Core loop
import requests

def answer(question):
    # 1. Find similar documents
    qvec = requests.post("http://localhost:11434/api/embeddings",
                         json={"model": "nomic-embed-text", "prompt": question}
                         ).json()["embedding"]
    hits = qdrant.search("firma", query_vector=qvec, limit=5)

    # 2. Build context
    context = "\n\n".join(h.payload["text"] for h in hits)
    sources = [h.payload["title"] for h in hits]

    # 3. Generate answer, ONLY from context
    prompt = f"""Answer the question ONLY from these documents.
Documents:
{context}

Question: {question}
If the answer is not in the documents: say so honestly."""

    r = requests.post("http://localhost:11434/api/generate",
                      json={"model": "qwen2.5:14b", "prompt": prompt, "stream": False})
    return r.json()["response"], sources

The key: “ONLY from these documents” in the prompt prevents hallucinations. The bot cites the archive instead of guessing.

Step 4: Connect Your Chat Frontend

The bot answers where your company already chats. Examples in the platform articles:

  • Mattermost (Article): WebSocket bot, /knowledge question..., self-hosted, GDPR-clean.
  • Nextcloud Talk (Article): Webhook bot, built right into your existing Nextcloud.
  • Simple web frontend: Open WebUI with RAG support, fastest path, no custom bot code needed.
  • Telegram (Article): For mobile too.

Step 5: Operations and Automatic Indexing

A cron job or Paperless webhook triggers ingest.py each time a document arrives. The archive grows, the bot learns automatically:

# Paperless: post-consume-script or simple cron
*/10 * * * * docker exec paperless python3 /scripts/ingest.py --new-only

Extensions

  • Per-department permissions: Separate Qdrant collections by department. Sales doesn’t see HR documents.
  • n8n variant: Same flow as an n8n workflow, visual, no Python.
  • Answer quality: Add a reranking model after embedding retrieval. See Reranking Models.
  • Source links: Bot posts direct Paperless links to documents, not just titles.
  • Multilingual: Foreign contracts. Choose a multilingual embedding model.

What You’ll Learn

  • RAG architecture in real production (not just theory)
  • OCR + archive + search as a pipeline
  • Hallucination prevention through “context only” prompts
  • Chat bots in business context (permissions, GDPR)

Further Reading

Key Takeaways:

  • Company document bot = Paperless (archive/OCR) + Qdrant (vector search) + Ollama (answers) + chat frontend.
  • “Answer only from documents” prompt prevents hallucinations. Core of the design.
  • Fully local means GDPR-clean for contracts and client data.
  • Indexing via cron or webhook: archive grows, bot learns with it.
  • About 2-3 hours to set up on existing hardware.

FAQ

Does this prevent hallucinations?

Largely yes. The prompt enforces “answer only from context” plus source attribution. No match found means the bot says “not found” instead of making something up. Reranking further improves hit quality.

How many documents can it handle?

Tens of thousands without trouble. Qdrant scales well. The real bottleneck is chunking strategy: split very long documents into sections for better matches.

What hardware do I need?

32 GB RAM is enough for 7B-14B answer models plus embeddings. For better quality, use 14B-30B models, then you need 64 GB or a 128 GB mini PC (MS-S1 Max). See the Homelab guide.

Per-department permissions?

Yes, separate Qdrant collections per department. The bot checks user role from chat (Mattermost or Nextcloud provides it) and searches only in allowed collections.

Can I do this without Python?

Yes, as an n8n workflow: webhook → Paperless search → Ollama → answer to chat. See the n8n RAG workflows article. Slightly less flexible, but no code required.

Sources and Further Reading

Back to Blog
Share:

Related Posts