Extending Paperless-ngx with AI
What this article covers
- How to extend Paperless-ngx with Ollama for AI capabilities.
- How automatic classification, tagging, and summarization work.
- How to use custom scripts and post-consume hooks.
- Practical examples for invoices, contracts, and correspondence.
- Best practices for accuracy, security, and performance.
Introduction: Understanding Paperless-ngx with AI
Paperless-ngx is a document management system. You scan documents, Paperless-ngx runs OCR, indexes them, and makes them searchable. AI takes it further: the model automatically classifies documents, assigns tags, creates summaries, and extracts metadata. Instead of sorting manually, AI does the work.
This article is for users who want to extend Paperless-ngx with AI capabilities. Foundational concepts are covered in Document Automation and Ollama.
Why would you need Paperless-ngx with AI?
Paperless-ngx handles OCR and full-text search, but classification relies on rules. With AI, the system understands content: “This is an invoice from Company X for €1,234, due March 15” instead of just “contains the word invoice”. The model classifies more accurately and extracts structured data.
Paperless-ngx with AI in a nutshell
Paperless-ngx processes documents (OCR, indexing). A post-consume script calls Ollama: the AI classifies, tags, and summarizes the document. The results are stored in Paperless-ngx.
The core idea: Paperless-ngx manages, AI understands.
Who is this article for?
- Paperless-ngx users who want AI features.
- Office workers who need automatic document processing.
- Self-hosters extending document management with AI.
- Organizations automating invoices and contracts.
Experience with Paperless-ngx and Ollama is helpful.
Key concepts
- Paperless-ngx - Document management. When useful: the foundation.
- Ollama - Local model server. When useful: the AI backend.
- Post-Consume Script - Script that runs after document import. When useful: for AI processing.
- OCR - Optical Character Recognition. When useful: for scanned documents.
- Tags - Keywords for categorization. When useful: organizing documents.
- Custom Fields - User-defined fields. When useful: storing AI-generated metadata.
- RAG - Retrieval-Augmented Generation. When useful: document Q&A.
Setup: Paperless-ngx + Ollama
Docker Compose
version: "3.8"
services:
paperless:
image: ghcr.io/paperless-ngx/paperless-ngx:latest
container_name: paperless
restart: unless-stopped
ports:
- "8000:8000"
volumes:
- paperless_data:/usr/src/paperless/data
- paperless_media:/usr/src/paperless/media
- ./scripts:/usr/src/paperless/scripts # Custom Scripts
environment:
- PAPERLESS_OCR_LANGUAGE=deu
- PAPERLESS_POST_CONSUME_SCRIPT=/usr/src/paperless/scripts/post_consume.py
- PAPERLESS_OLLAMA_URL=http://ollama:11434
networks:
- paperless-network
ollama:
image: ollama/ollama:latest
container_name: ollama
restart: unless-stopped
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
networks:
- paperless-network
volumes:
paperless_data:
paperless_media:
ollama_data:
networks:
paperless-network:
driver: bridge
Post-Consume Script
#!/usr/bin/env python3
# /usr/src/paperless/scripts/post_consume.py
# Runs after every document import
import os
import sys
import json
import requests
OLLAMA_URL = os.environ.get("PAPERLESS_OLLAMA_URL", "http://ollama:11434")
def call_ollama(prompt, model="llama3.1", format=None):
"""Call Ollama"""
body = {
"model": model,
"messages": [{"role": "user", "content": prompt}],
"stream": False
}
if format:
body["format"] = format
response = requests.post(
f"{OLLAMA_URL}/api/chat",
json=body,
timeout=120
)
return response.json()["message"]["content"]
def main():
# Paperless-ngx passes document info as arguments
doc_id = sys.argv[1] if len(sys.argv) > 1 else None
doc_file = sys.argv[2] if len(sys.argv) > 2 else None
# Read document text (already OCR-processed)
text_file = doc_file.replace(".pdf", ".txt")
if os.path.exists(text_file):
with open(text_file) as f:
text = f.read()
else:
text = ""
if not text:
return
# 1. Classify
category = call_ollama(f"""Classify the document:
- invoice: Invoice, Receipt
- contract: Contract, Agreement
- report: Report, Statement
- correspondence: Letter, Email
- other: Everything else
Document: {text[:3000]}
Reply with only the category.""")
# 2. Generate tags
tags = call_ollama(f"""Assign 3-5 tags for this document.
Document: {text[:2000]}
Reply as JSON array: ["tag1", "tag2", ...]""", format="json")
# 3. Create summary
summary = call_ollama(f"""Summarize the document in 3 sentences.
Document: {text[:4000]}""")
# 4. Extract metadata
metadata = call_ollama(f"""Extract:
- date: Document date
- sender: Who created it
- subject: What it is about
- amount: If invoice, the amount
Document: {text[:3000]}
Reply as JSON.""", format="json")
# Output for Paperless-ngx
result = {
"category": category.strip().lower(),
"tags": json.loads(tags),
"summary": summary,
"metadata": json.loads(metadata)
}
print(json.dumps(result))
if __name__ == "__main__":
main()
Practical example 1: Automating invoice processing
def process_invoice(text):
"""Process invoice"""
# Extract structured data
data = call_ollama(f"""Extract from the invoice:
- invoice_number
- date
- amount (numbers only)
- currency
- sender
- due_date
Invoice: {text}
Reply as JSON.""", format="json")
invoice = json.loads(data)
# Save to database or as custom fields
save_invoice_metadata(doc_id, invoice)
# Move to folder
move_to_folder(doc_id, "invoices")
return invoice
Practical example 2: Analyzing contracts
def analyze_contract(text):
"""Analyze contract"""
analysis = call_ollama(f"""Analyze the contract:
1. Type (Lease, Purchase, Services, ...)
2. Parties
3. Duration
4. Termination clause
5. Special provisions
6. Risks
Contract: {text[:6000]}
Reply as JSON.""", format="json")
result = json.loads(analysis)
# Add risks as comment
if result.get("risks"):
add_comment(doc_id, f"Risks: {result['risks']}")
return result
Practical example 3: Summarizing correspondence
def summarize_correspondence(text):
"""Summarize correspondence"""
summary = call_ollama(f"""Summarize the correspondence:
- Who is writing to whom?
- What is it about?
- What is being requested or offered?
- Are there deadlines?
Text: {text[:4000]}""")
# Save as custom field
set_custom_field(doc_id, "summary", summary)
return summary
Workflow: Complete Pipeline
Document scanned/uploaded
│
▼
Paperless-ngx OCR
│
▼
Post-Consume Script
│
├─► Ollama: Classify
├─► Ollama: Generate tags
├─► Ollama: Summarize
├─► Ollama: Extract metadata
│
▼
Paperless-ngx stores:
- Tags
- Custom Fields
- Comments
│
▼
Optional: Move to folder
Optional: Alert on critical documents
RAG for Documents
def setup_rag(documents):
"""Index all documents for semantic search"""
for doc in documents:
# Split text into chunks
chunks = split_into_chunks(doc.text, 500)
for i, chunk in enumerate(chunks):
# Create embedding
embedding = get_embedding(chunk)
# Store in Qdrant
qdrant.upsert(
collection="dokumente",
points=[{
"id": f"{doc.id}_{i}",
"vector": embedding,
"payload": {
"doc_id": doc.id,
"text": chunk,
"title": doc.title
}
}]
)
Now you can ask questions about your documents:
- “What does the lease agreement say about termination?”
- “Which invoices are over 1000 €?”
See Local RAG.
Security Considerations
- Confidential documents: Documents may contain sensitive data. Use local AI, not cloud services. See Data Protection.
- Prompt injection: Documents can contain injections. See Prompt Injection.
- Access control: Paperless-ngx supports user and group permissions. Use them.
- Audit logging: Log all AI processing. See Audit Logging.
- Backups: Back up documents and database regularly. See Backups.
Common Pitfalls
- OCR errors: Scanned documents contain OCR mistakes. AI results can be inaccurate as a result.
- Context length: Large documents need to be chunked. See Context Length.
- Misclassification: AI can classify incorrectly. Use confidence scores or manual review.
- Performance: Large documents plus many AI calls equals slow processing. Process asynchronously.
- Script failures: If the post-consume script fails, the document won’t be processed. Add logging.
Further Reading
- Document Automation - Automate documents with AI.
- Local RAG - Make documents searchable.
- Ollama - Local model server.
- PDF and OCR - Prepare documents.
- Backups - Data protection.
- Prompt Injection - Security.
Key Takeaways:
- Paperless-ngx plus Ollama: OCR plus AI understanding.
- Post-Consume Script for automatic classification, tagging, summarization.
- Structured data extraction for invoices, contracts, correspondence.
- RAG for semantic document search.
- Security: validation, access control, local processing.
FAQ
How do I connect Paperless-ngx with Ollama?
What can AI automate?
Do I still need OCR?
How accurate is the classification?
Can I search through documents?
How fast is processing?
Are my documents safe?
Which formats are supported?
Sources and Further Reading
- Paperless-ngx - Documentation.
- Paperless-ngx Scripts - Post-Consume Scripts.
- Ollama - Local model server.
- Tesseract OCR - OCR engine.


