Project: Building a Company Document Bot: Paperless-ngx + Ollama + RAG in Your Own Chat
What This Project Does
Employees ask in chat: “What’s the notice period in the contract with Customer X?” or “Where’s the manual for the machine?” The bot searches the company archive and responds with sources. Completely local: no documents leave the office.
The Stack:
Paper/Email/PDF
└── Paperless-ngx (Scan + OCR + Archive)
└── RAG Pipeline (Embeddings + Qdrant)
└── Ollama (Answer Model)
└── Chat Frontend (Mattermost / Nextcloud Talk / Web)
Prerequisites: A mini PC or server (32 GB RAM is enough), the tools from the Software Stack article, about 2-3 hours for setup.
Why This Architecture
- Paperless-ngx is the archive backbone: OCR makes scans searchable, tags and correspondents organize content. See Paperless-ngx.
- RAG instead of training: Documents aren’t baked into the model; they’re searched via embeddings. New documents appear instantly, no retraining needed.
- Chat as the frontend: Employees ask where they already are: Mattermost or Nextcloud Talk. No new tool to learn.
- GDPR-compliant: Everything on-premise. Client and contract data stays internal.
Step 1: Paperless-ngx as Archive
# docker-compose.yml (core)
services:
paperless:
image: ghcr.io/paperless-ngx/paperless-ngx:latest
ports: ["127.0.0.1:8000:8000"]
environment:
- PAPERLESS_OCR_LANGUAGE=deu
- PAPERLESS_TIME_ZONE=Europe/Berlin
volumes:
- paperless_data:/usr/src/paperless/data
- ./consume:/usr/src/paperless/consume # Drop files here → OCR + Archive
volumes:
paperless_data:
Drop documents into ./consume (or point your scanner directly there): Paperless handles OCR and archival. See Paperless-ngx Integration.
Step 2: RAG Pipeline: Indexing Documents for Search
# ingest.py, Index Paperless documents into Qdrant
import requests, uuid
from qdrant_client import QdrantClient
from qdrant_client.models import PointStruct
qdrant = QdrantClient("localhost", port=6333)
def embed(text):
r = requests.post("http://localhost:11434/api/embeddings",
json={"model": "nomic-embed-text", "prompt": text})
return r.json()["embedding"]
# Paperless API: fetch documents
docs = requests.get("http://localhost:8000/api/documents/",
headers={"Authorization": "Token YOUR_TOKEN"}).json()["results"]
for doc in docs:
text = doc.get("content", "")[:8000] # OCR text
vec = embed(text)
qdrant.upsert("firma", [PointStruct(
id=str(uuid.uuid4()), vector=vec,
payload={"title": doc["title"], "id": doc["id"], "text": text[:2000]}
)])
Use nomic-embed-text as the embedding model (ollama pull nomic-embed-text), Qdrant as the vector database. See Local RAG.
Step 3: The Bot: Question → Retrieval → Answer
# bot.py, Core loop
import requests
def answer(question):
# 1. Find similar documents
qvec = requests.post("http://localhost:11434/api/embeddings",
json={"model": "nomic-embed-text", "prompt": question}
).json()["embedding"]
hits = qdrant.search("firma", query_vector=qvec, limit=5)
# 2. Build context
context = "\n\n".join(h.payload["text"] for h in hits)
sources = [h.payload["title"] for h in hits]
# 3. Generate answer, ONLY from context
prompt = f"""Answer the question ONLY from these documents.
Documents:
{context}
Question: {question}
If the answer is not in the documents: say so honestly."""
r = requests.post("http://localhost:11434/api/generate",
json={"model": "qwen2.5:14b", "prompt": prompt, "stream": False})
return r.json()["response"], sources
The key: “ONLY from these documents” in the prompt prevents hallucinations. The bot cites the archive instead of guessing.
Step 4: Connect Your Chat Frontend
The bot answers where your company already chats. Examples in the platform articles:
- Mattermost (Article): WebSocket bot,
/knowledge question..., self-hosted, GDPR-clean. - Nextcloud Talk (Article): Webhook bot, built right into your existing Nextcloud.
- Simple web frontend: Open WebUI with RAG support, fastest path, no custom bot code needed.
- Telegram (Article): For mobile too.
Step 5: Operations and Automatic Indexing
A cron job or Paperless webhook triggers ingest.py each time a document arrives. The archive grows, the bot learns automatically:
# Paperless: post-consume-script or simple cron
*/10 * * * * docker exec paperless python3 /scripts/ingest.py --new-only
Extensions
- Per-department permissions: Separate Qdrant collections by department. Sales doesn’t see HR documents.
- n8n variant: Same flow as an n8n workflow, visual, no Python.
- Answer quality: Add a reranking model after embedding retrieval. See Reranking Models.
- Source links: Bot posts direct Paperless links to documents, not just titles.
- Multilingual: Foreign contracts. Choose a multilingual embedding model.
What You’ll Learn
- RAG architecture in real production (not just theory)
- OCR + archive + search as a pipeline
- Hallucination prevention through “context only” prompts
- Chat bots in business context (permissions, GDPR)
Further Reading
- IRC-Coding.de: Programming tutorials on RAG, bot code.
- Local RAG: The technology in depth.
- Paperless-ngx: Archive setup.
- Mattermost Bots: Chat frontend.
- Internal Knowledge Bot: Related scenario.
- Homelab: The hardware.
Key Takeaways:
- Company document bot = Paperless (archive/OCR) + Qdrant (vector search) + Ollama (answers) + chat frontend.
- “Answer only from documents” prompt prevents hallucinations. Core of the design.
- Fully local means GDPR-clean for contracts and client data.
- Indexing via cron or webhook: archive grows, bot learns with it.
- About 2-3 hours to set up on existing hardware.
FAQ
Does this prevent hallucinations?
How many documents can it handle?
What hardware do I need?
Per-department permissions?
Can I do this without Python?
Sources and Further Reading
- Paperless-ngx: GitHub.
- Qdrant: Vector database.
- IRC-Coding.de: Programming tutorials.


