Skip to content
BotServBotServ
Paperless-ngxAIDocument ManagementOCRClassification

Extend Paperless-ngx with AI

Enhance Paperless-ngx with AI: automatic classification, OCR, tagging, summarization and practical examples.

S

schutzgeist

6 min read
Extend Paperless-ngx with AI

Extending Paperless-ngx with AI

What this article covers

  • How to extend Paperless-ngx with Ollama for AI capabilities.
  • How automatic classification, tagging, and summarization work.
  • How to use custom scripts and post-consume hooks.
  • Practical examples for invoices, contracts, and correspondence.
  • Best practices for accuracy, security, and performance.

Introduction: Understanding Paperless-ngx with AI

Paperless-ngx is a document management system. You scan documents, Paperless-ngx runs OCR, indexes them, and makes them searchable. AI takes it further: the model automatically classifies documents, assigns tags, creates summaries, and extracts metadata. Instead of sorting manually, AI does the work.

This article is for users who want to extend Paperless-ngx with AI capabilities. Foundational concepts are covered in Document Automation and Ollama.

Why would you need Paperless-ngx with AI?

Paperless-ngx handles OCR and full-text search, but classification relies on rules. With AI, the system understands content: “This is an invoice from Company X for €1,234, due March 15” instead of just “contains the word invoice”. The model classifies more accurately and extracts structured data.

Paperless-ngx with AI in a nutshell

Paperless-ngx processes documents (OCR, indexing). A post-consume script calls Ollama: the AI classifies, tags, and summarizes the document. The results are stored in Paperless-ngx.

The core idea: Paperless-ngx manages, AI understands.

Who is this article for?

  • Paperless-ngx users who want AI features.
  • Office workers who need automatic document processing.
  • Self-hosters extending document management with AI.
  • Organizations automating invoices and contracts.

Experience with Paperless-ngx and Ollama is helpful.

Key concepts

  • Paperless-ngx - Document management. When useful: the foundation.
  • Ollama - Local model server. When useful: the AI backend.
  • Post-Consume Script - Script that runs after document import. When useful: for AI processing.
  • OCR - Optical Character Recognition. When useful: for scanned documents.
  • Tags - Keywords for categorization. When useful: organizing documents.
  • Custom Fields - User-defined fields. When useful: storing AI-generated metadata.
  • RAG - Retrieval-Augmented Generation. When useful: document Q&A.

Setup: Paperless-ngx + Ollama

Docker Compose

version: "3.8"

services:
  paperless:
    image: ghcr.io/paperless-ngx/paperless-ngx:latest
    container_name: paperless
    restart: unless-stopped
    ports:
      - "8000:8000"
    volumes:
      - paperless_data:/usr/src/paperless/data
      - paperless_media:/usr/src/paperless/media
      - ./scripts:/usr/src/paperless/scripts  # Custom Scripts
    environment:
      - PAPERLESS_OCR_LANGUAGE=deu
      - PAPERLESS_POST_CONSUME_SCRIPT=/usr/src/paperless/scripts/post_consume.py
      - PAPERLESS_OLLAMA_URL=http://ollama:11434
    networks:
      - paperless-network

  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "11434:11434"
    volumes:
      - ollama_data:/root/.ollama
    networks:
      - paperless-network

volumes:
  paperless_data:
  paperless_media:
  ollama_data:

networks:
  paperless-network:
    driver: bridge

Post-Consume Script

#!/usr/bin/env python3
# /usr/src/paperless/scripts/post_consume.py
# Runs after every document import

import os
import sys
import json
import requests

OLLAMA_URL = os.environ.get("PAPERLESS_OLLAMA_URL", "http://ollama:11434")

def call_ollama(prompt, model="llama3.1", format=None):
    """Call Ollama"""
    body = {
        "model": model,
        "messages": [{"role": "user", "content": prompt}],
        "stream": False
    }
    if format:
        body["format"] = format

    response = requests.post(
        f"{OLLAMA_URL}/api/chat",
        json=body,
        timeout=120
    )
    return response.json()["message"]["content"]

def main():
    # Paperless-ngx passes document info as arguments
    doc_id = sys.argv[1] if len(sys.argv) > 1 else None
    doc_file = sys.argv[2] if len(sys.argv) > 2 else None

    # Read document text (already OCR-processed)
    text_file = doc_file.replace(".pdf", ".txt")
    if os.path.exists(text_file):
        with open(text_file) as f:
            text = f.read()
    else:
        text = ""

    if not text:
        return

    # 1. Classify
    category = call_ollama(f"""Classify the document:
- invoice: Invoice, Receipt
- contract: Contract, Agreement
- report: Report, Statement
- correspondence: Letter, Email
- other: Everything else

Document: {text[:3000]}

Reply with only the category.""")

    # 2. Generate tags
    tags = call_ollama(f"""Assign 3-5 tags for this document.
Document: {text[:2000]}
Reply as JSON array: ["tag1", "tag2", ...]""", format="json")

    # 3. Create summary
    summary = call_ollama(f"""Summarize the document in 3 sentences.
Document: {text[:4000]}""")

    # 4. Extract metadata
    metadata = call_ollama(f"""Extract:
- date: Document date
- sender: Who created it
- subject: What it is about
- amount: If invoice, the amount

Document: {text[:3000]}
Reply as JSON.""", format="json")

    # Output for Paperless-ngx
    result = {
        "category": category.strip().lower(),
        "tags": json.loads(tags),
        "summary": summary,
        "metadata": json.loads(metadata)
    }
    print(json.dumps(result))

if __name__ == "__main__":
    main()

Practical example 1: Automating invoice processing

def process_invoice(text):
    """Process invoice"""
    # Extract structured data
    data = call_ollama(f"""Extract from the invoice:
- invoice_number
- date
- amount (numbers only)
- currency
- sender
- due_date

Invoice: {text}
Reply as JSON.""", format="json")

    invoice = json.loads(data)

    # Save to database or as custom fields
    save_invoice_metadata(doc_id, invoice)

    # Move to folder
    move_to_folder(doc_id, "invoices")

    return invoice

Practical example 2: Analyzing contracts

def analyze_contract(text):
    """Analyze contract"""
    analysis = call_ollama(f"""Analyze the contract:
1. Type (Lease, Purchase, Services, ...)
2. Parties
3. Duration
4. Termination clause
5. Special provisions
6. Risks

Contract: {text[:6000]}

Reply as JSON.""", format="json")

    result = json.loads(analysis)

    # Add risks as comment
    if result.get("risks"):
        add_comment(doc_id, f"Risks: {result['risks']}")

    return result

Practical example 3: Summarizing correspondence

def summarize_correspondence(text):
    """Summarize correspondence"""
    summary = call_ollama(f"""Summarize the correspondence:
- Who is writing to whom?
- What is it about?
- What is being requested or offered?
- Are there deadlines?

Text: {text[:4000]}""")

    # Save as custom field
    set_custom_field(doc_id, "summary", summary)

    return summary

Workflow: Complete Pipeline

Document scanned/uploaded
    │
    ▼
Paperless-ngx OCR
    │
    ▼
Post-Consume Script
    │
    ├─► Ollama: Classify
    ├─► Ollama: Generate tags
    ├─► Ollama: Summarize
    ├─► Ollama: Extract metadata
    │
    ▼
Paperless-ngx stores:
  - Tags
  - Custom Fields
  - Comments
    │
    ▼
Optional: Move to folder
Optional: Alert on critical documents

RAG for Documents

def setup_rag(documents):
    """Index all documents for semantic search"""
    for doc in documents:
        # Split text into chunks
        chunks = split_into_chunks(doc.text, 500)

        for i, chunk in enumerate(chunks):
            # Create embedding
            embedding = get_embedding(chunk)

            # Store in Qdrant
            qdrant.upsert(
                collection="dokumente",
                points=[{
                    "id": f"{doc.id}_{i}",
                    "vector": embedding,
                    "payload": {
                        "doc_id": doc.id,
                        "text": chunk,
                        "title": doc.title
                    }
                }]
            )

Now you can ask questions about your documents:

  • “What does the lease agreement say about termination?”
  • “Which invoices are over 1000 €?”

See Local RAG.

Security Considerations

  • Confidential documents: Documents may contain sensitive data. Use local AI, not cloud services. See Data Protection.
  • Prompt injection: Documents can contain injections. See Prompt Injection.
  • Access control: Paperless-ngx supports user and group permissions. Use them.
  • Audit logging: Log all AI processing. See Audit Logging.
  • Backups: Back up documents and database regularly. See Backups.

Common Pitfalls

  • OCR errors: Scanned documents contain OCR mistakes. AI results can be inaccurate as a result.
  • Context length: Large documents need to be chunked. See Context Length.
  • Misclassification: AI can classify incorrectly. Use confidence scores or manual review.
  • Performance: Large documents plus many AI calls equals slow processing. Process asynchronously.
  • Script failures: If the post-consume script fails, the document won’t be processed. Add logging.

Further Reading

Key Takeaways:

  • Paperless-ngx plus Ollama: OCR plus AI understanding.
  • Post-Consume Script for automatic classification, tagging, summarization.
  • Structured data extraction for invoices, contracts, correspondence.
  • RAG for semantic document search.
  • Security: validation, access control, local processing.

FAQ

How do I connect Paperless-ngx with Ollama?

Via a post-consume script. Paperless-ngx runs the script after each import. The script sends the OCR text to Ollama and processes the results.

What can AI automate?

Classify documents (invoice, contract, etc.), assign tags, generate summaries, extract metadata (date, amount, sender), identify risks.

Do I still need OCR?

Yes. Paperless-ngx handles OCR. The AI works with the extracted text. For scanned documents, OCR is step one and AI is step two.

How accurate is the classification?

Very good for standard documents like invoices and contracts. For unusual documents, the AI can be wrong. For critical documents, use manual review.

Can I search through documents?

Yes, with RAG. All documents are chunked, embedded, and stored in a vector database. Then you can ask questions: “What does the contract say about termination?”

How fast is processing?

OCR: seconds. AI processing: 10-60 seconds depending on document size and model. For many documents, process asynchronously.

Are my documents safe?

Yes, if you run Ollama locally. All data stays on your server. With cloud APIs, documents go to the cloud, which is unacceptable for confidential documents without local AI.

Which formats are supported?

PDF (with and without text layer), images (PNG, JPG via OCR), Word, plain text. Converters exist for other formats.

Sources and Further Reading

Back to Blog
Share:

Related Posts