Skip to content
BotServBotServ
Document AutomationAI AgentsPDFExtractionClassification

Document Automation with AI Agents

Automate document processing with AI agents. Extract, classify, summarize PDFs with practical examples.

S

schutzgeist

6 min read
Document Automation with AI Agents

Document Automation with AI Agents

What this article covers

  • How to automatically process documents using AI agents.
  • How PDF extraction, classification, summarization, and tagging work.
  • How to integrate Paperless-ngx, Nextcloud, and custom scripts with AI.
  • Practical examples for invoices, contracts, reports, and email attachments.
  • Best practices for accuracy, security, and performance.

Introduction: document automation with AI agents explained

Document automation means AI agents handle the document processing workflow: reading PDFs, extracting content, classifying documents, creating summaries, and applying tags. Instead of manually sorting and reading documents, an agent does it automatically. An incoming PDF gets analyzed, its content extracted, the document classified, and tags applied.

This article is for users who want to automate document processing. You should understand what AI agents are and how Ollama works. For Python programming fundamentals, see IRC-Coding.de.

Why do you need document automation?

Imagine receiving 20 PDFs daily: invoices, contracts, reports, quotes. Each one requires reading, categorizing, and filing. That’s hours of work. With AI agents, the PDF is automatically analyzed, content extracted, the document classified, a summary generated, and tags applied. You see only the summary and tags.

How document automation with AI agents works

Document automation uses AI agents to process documents. The agent reads the document (PDF, Word, text), extracts content, classifies the document, creates a summary, and assigns tags. All automatically, with no human intervention.

The core concept: documents come in, processed documents come out.

Who is this article for?

  • Office workers processing many documents daily.
  • Accountants who want to classify invoices automatically.
  • Lawyers who want to analyze contracts automatically.
  • Self-hosters extending Paperless-ngx or Nextcloud with AI.

Background knowledge of AI and Python is helpful.

Key terminology

  • AI agents - Programs with tools. Useful for: intelligent document processing.
  • PDF extraction - Reading text from PDFs. Useful for: processing PDF files.
  • Classification - Assigning a document to a category. Useful for: automatic sorting.
  • Summarization - Extracting key points. Useful for: quick understanding.
  • Tagging - Assigning keywords. Useful for: search and filtering.
  • OCR - Optical Character Recognition. Useful for: scanned documents.
  • Ollama - Local model server. Useful for: the AI backend.
  • Paperless-ngx - Document management. Useful for: document storage and retrieval.
  • Function Calling - Tool use. Useful for: structured extraction.

Document pipeline

A typical document pipeline flows through these stages:

  1. Intake: Document arrives (email, upload, scan).
  2. Extraction: Read text from document (PDF, OCR).
  3. Analysis: AI analyzes content.
  4. Classification: AI assigns a category.
  5. Summarization: AI creates a summary.
  6. Tagging: AI assigns tags.
  7. Storage: Document is filed.

PDF extraction

import PyPDF2
import pdfplumber

def extract_pdf_pypdf2(file_path):
    """Simple extraction with PyPDF2"""
    with open(file_path, "rb") as f:
        reader = PyPDF2.PdfReader(f)
        text = ""
        for page in reader.pages:
            text += page.extract_text()
    return text

def extract_pdf_pdfplumber(file_path):
    """Better extraction with pdfplumber (tables, layout)"""
    text = ""
    with pdfplumber.open(file_path) as pdf:
        for page in pdf.pages:
            text += page.extract_text() + "\n"
    return text

OCR for scanned documents

import pytesseract
from PIL import Image

def extract_ocr(image_path):
    """OCR for scanned documents"""
    image = Image.open(image_path)
    text = pytesseract.image_to_string(image, lang="deu")
    return text

Document classification with AI

def classify_document(text):
    """Classify document"""
    response = call_ollama([
        {"role": "system", "content": """Classify the document into one category:
- invoice: Invoice, payment request
- contract: Contract, agreement
- report: Report, analysis
- quote: Quote, proposal
- other: Everything else

Reply with only the category."""},
        {"role": "user", "content": text[:3000]}
    ])
    return response["message"]["content"].strip().lower()

Document summarization with AI

def summarize_document(text, max_sentences=5):
    """Summarize document"""
    response = call_ollama([
        {"role": "system", "content": f"Summarize the document in {max_sentences} sentences."},
        {"role": "user", "content": text[:8000]}
    ])
    return response["message"]["content"]

Tagging with AI

def tag_document(text):
    """Generate tags for document"""
    response = call_ollama([
        {"role": "system", "content": "Assign 3-5 tags for the document. Reply as a JSON array."},
        {"role": "user", "content": text[:3000]}
    ], format="json")
    return json.loads(response["message"]["content"])

Complete pipeline

def process_document(file_path):
    """Complete document processing pipeline"""
    # 1. Extract text
    if file_path.endswith(".pdf"):
        text = extract_pdf_pdfplumber(file_path)
    elif file_path.endswith((".png", ".jpg")):
        text = extract_ocr(file_path)
    else:
        with open(file_path) as f:
            text = f.read()

    # 2. Classify
    category = classify_document(text)

    # 3. Summarize
    summary = summarize_document(text)

    # 4. Generate tags
    tags = tag_document(text)

    # 5. Extract metadata (date, sender, subject)
    metadata = extract_metadata(text)

    return {
        "file": file_path,
        "category": category,
        "summary": summary,
        "tags": tags,
        "metadata": metadata
    }

Practical example 1: Processing invoices automatically

def process_invoice(pdf_path):
    """Process invoice"""
    text = extract_pdf_pdfplumber(pdf_path)

    # Extract structured data
    invoice_data = extract_invoice_data(text)

    # Save to database
    save_invoice(invoice_data)

    # Move to folder
    move_to_folder(pdf_path, "invoices")

    return invoice_data

def extract_invoice_data(text):
    """Extract invoice data"""
    response = call_ollama([
        {"role": "system", "content": """Extract from the invoice:
- Invoice number
- Date
- Amount
- Currency
- Sender
- Recipient

Reply as JSON."""},
        {"role": "user", "content": text}
    ], format="json")
    return json.loads(response["message"]["content"])

Practical Example 2: Analyzing Contracts

def analyze_contract(pdf_path):
    """Analyze contract"""
    text = extract_pdf_pdfplumber(pdf_path)

    # Identify risks
    risks = identify_risks(text)

    # Extract key clauses
    clauses = extract_clauses(text)

    # Generate summary
    summary = summarize_document(text, max_sentences=10)

    return {
        "risks": risks,
        "clauses": clauses,
        "summary": summary
    }

def identify_risks(text):
    """Identify risks in contract"""
    response = call_ollama([
        {"role": "system", "content": "Identify risks in the contract. List them."},
        {"role": "user", "content": text[:8000]}
    ])
    return response["message"]["content"]

Practical Example 3: Paperless-ngx with AI

# Paperless-ngx Custom Script
# Called each time a new document is processed

def on_document_processed(document):
    """Called by Paperless-ngx"""
    text = document.content

    # Classify
    category = classify_document(text)
    document.tags.add(category)

    # Summarize
    summary = summarize_document(text)
    document.custom_fields["summary"] = summary

    # Tag
    tags = tag_document(text)
    for tag in tags:
        document.tags.add(tag)

Security Considerations

  • Confidential Documents: Documents may contain sensitive data. Use local AI, not cloud services. See Offline AI.
  • Validate Input: Documents can contain prompt injection attacks. See Prompt Injection.
  • Validate Output: AI can make mistakes. Verify critical extractions.
  • Audit Logging: Log all document processing. See Logging.
  • Access Control: Not everyone should see all documents. See Access Protection.

Common Pitfalls

  • OCR Errors: Scanned documents contain OCR artifacts. Verify critical data.
  • Context Length: Large documents need to be chunked. Use Context Length.
  • Misclassification: AI can classify incorrectly. Use confidence scores or human review.
  • Extraction Errors: AI can extract data incorrectly. Validate critical fields.
  • Performance: Large documents take time to process. Use asynchronous processing.
  • Layout Issues: Complex layouts (tables, columns) are hard to extract.

Further Reading

Key Takeaways:

  • Document automation: extraction, classification, summarization, tagging.
  • PDF extraction with PyPDF2 or pdfplumber, OCR for scans.
  • AI classifies, summarizes, and assigns tags.
  • Paperless-ngx can be extended with custom scripts.
  • Security: validation, audit logging, access control.

FAQ

What is document automation with AI?

AI agents process documents automatically: read PDFs, extract content, classify, summarize, and assign tags. Instead of manually sorting documents, an agent does it for you.

How do I extract text from PDFs?

Use PyPDF2 for basic extraction or pdfplumber for better results (tables, layout). For scanned documents, use OCR (pytesseract).

How do I classify documents?

Use an AI model (Ollama) with a prompt that defines categories. The model reads the text and returns the category.

How do I create summaries?

Send the document text to the model with instructions to summarize it in X sentences. For long documents, chunk them and summarize each chunk.

How do I integrate Paperless-ngx?

Using custom scripts. Paperless-ngx calls your script each time a new document arrives. The script classifies, summarizes, and assigns tags.

How accurate is the AI?

For standard tasks (classification, summarization) very good. For critical extractions (invoice data), you should validate. AI can make mistakes.

What is OCR?

Optical Character Recognition. For scanned documents that don’t have a text layer. pytesseract is a popular OCR library.

Is this safe for confidential documents?

Yes, if you use local AI (Ollama). All data stays local. With cloud APIs, documents go to the cloud. For confidential data, local AI is essential.

What does this cost?

With Ollama locally: only hardware costs. No API fees. A machine with an RTX 4070 is sufficient for most document tasks.

What formats can I process?

PDF (PyPDF2, pdfplumber), images (OCR), Word (python-docx), text. For other formats, there are libraries or converters.

Sources and Further Reading

Back to Blog
Share:

Related Posts