Skip to content
BotServBotServ
Document AnalysisOCRSummarizationClassificationMetadata

Document Analysis with Local AI

Analyze documents with local AI. Text extraction, summarization, metadata, classification and practical examples.

S

schutzgeist

4 min read
Document Analysis with Local AI

Document Analysis with Local AI

What this article covers

  • How to analyze documents using local AI.
  • How text extraction, summarization, and metadata extraction work.
  • How to classify and tag documents automatically.
  • Practical examples for invoices, contracts, reports, and correspondence.
  • Best practices for accuracy, privacy, and performance.

Introduction: Document analysis explained

Document analysis means the AI reads a document and understands its content. It can summarize, classify, extract metadata, and answer questions. Using Ollama locally keeps all documents private.

This article is for anyone who wants to analyze documents with AI. You’ll find foundational material in Documents and PDF and Ollama.

Why do you need document analysis?

Imagine you have 500 PDFs: invoices, contracts, reports. Instead of reading each one manually, the AI analyzes: “This is an invoice from Company X for 1,234 EUR, due on March 15. Risk: no payment terms specified.” The AI transforms unstructured documents into structured data.

Document analysis at a glance

Document → Extract text (OCR if needed) → AI analyzes → Structured data (category, tags, summary, metadata). All handled locally by Ollama.

The core concept: turn unstructured text into structured data.

Who is this article for?

  • Office workers who need to process documents automatically.
  • Businesses analyzing invoices and contracts.
  • Developers building document pipelines.
  • Self-hosters running document analysis locally.

Key terms

  • Ollama - Local model server. Useful for: the AI backend.
  • OCR - Text from images. Useful for: scanned documents.
  • Classification - Categorizing documents. Useful for: organizing files.
  • Metadata - Structured data. Useful for: databases.
  • RAG - Document Q&A. Useful for: asking questions.
  • Vision models - For images. Useful for: scanning documents.

The pipeline

Document (PDF/Word/Image)
    │
    ▼
Text Extraction
    ├─ Text present? → Use directly
    └─ No text? → OCR (Tesseract/Vision model)
    │
    ▼
AI Analysis (Ollama)
    ├─ Classify: Invoice? Contract? Report?
    ├─ Generate tags: topics, categories
    ├─ Summarize: key content
    └─ Extract metadata: date, sender, amount, ...
    │
    ▼
Structured Data
    ├─ Store in database
    ├─ Move to folder
    └─ Export as JSON

Practical example 1: Analyzing invoices

import requests
import json

def analyze_invoice(text):
    """Analyze invoice"""
    response = requests.post("http://ollama:11434/api/chat", json={
        "model": "llama3.1",
        "messages": [
            {"role": "system", "content": "Extract invoice data as JSON: invoice_number, date, amount, currency, sender, due_date, line_items (array)."},
            {"role": "user", "content": f"Invoice:\n{text[:4000]}"}
        ],
        "stream": False,
        "format": "json"
    })
    return json.loads(response.json()["message"]["content"])

# Example
invoice_data = analyze_invoice(invoice_text)
# Output: {"invoice_number": "RE-2024-1234", "amount": 1234.56, "currency": "EUR", ...}

Practical example 2: Analyzing contracts

def analyze_contract(text):
    """Analyze contract"""
    response = requests.post("http://ollama:11434/api/chat", json={
        "model": "llama3.1",
        "messages": [
            {"role": "system", "content": "Analyze the contract as JSON: contract_type, parties, duration, termination_notice, clauses (array), risks (array)."},
            {"role": "user", "content": f"Contract:\n{text[:6000]}"}
        ],
        "stream": False,
        "format": "json"
    })
    return json.loads(response.json()["message"]["content"])

Practical example 3: Classifying documents

def classify_document(text):
    """Classify document"""
    categories = [
        "invoice", "contract", "report", "correspondence",
        "minutes", "manual", "other"
    ]

    response = requests.post("http://ollama:11434/api/chat", json={
        "model": "llama3.1",
        "messages": [
            {"role": "system", "content": f"Classify the document into one of these categories: {', '.join(categories)}. Respond with only the category name."},
            {"role": "user", "content": f"Document:\n{text[:2000]}"}
        ],
        "stream": False
    })
    return response.json()["message"]["content"].strip().lower()

Practical example 4: Generating tags

def generate_tags(text, max_tags=5):
    """Generate tags for document"""
    response = requests.post("http://ollama:11434/api/chat", json={
        "model": "llama3.1",
        "messages": [
            {"role": "system", "content": f"Assign {max_tags} relevant tags for the document. Respond as a JSON array."},
            {"role": "user", "content": f"Document:\n{text[:2000]}"}
        ],
        "stream": False,
        "format": "json"
    })
    return json.loads(response.json()["message"]["content"])

Practical example 5: Complete pipeline

import os
from pathlib import Path

def process_document(filepath):
    """Complete document analysis"""
    # 1. Extract text
    text = extract_text(filepath)  # PDF, DOCX, TXT

    # 2. Classify
    category = classify_document(text)

    # 3. Generate tags
    tags = generate_tags(text)

    # 4. Summarize
    summary = summarize(text)

    # 5. Extract metadata
    metadata = extract_metadata(text)

    # 6. Store results
    result = {
        "file": filepath,
        "category": category,
        "tags": tags,
        "summary": summary,
        "metadata": metadata
    }

    save_to_database(result)
    move_to_folder(filepath, category)

    return result

Security considerations

  • Confidential documents: All data stays local with Ollama. Cloud APIs send documents outbound. See Privacy.
  • Prompt injection: Documents can contain injections. See Prompt Injection.
  • OCR errors: Scanned documents have OCR errors. AI results can be incorrect as a result.
  • Validation: AI extractions should be validated for critical data.

Common pitfalls

  • OCR errors: Poor scans mean poor extraction. Quality matters.
  • Context length: Large documents need to be chunked. See Context length.
  • Misclassification: AI can classify incorrectly. For critical documents, human review is needed.
  • Too many tags: More than 5-7 tags become unwieldy. Set a limit.
  • Performance: Large documents plus multiple AI calls slow things down. Process asynchronously.

Further reading

Key takeaways:

  • Document analysis: extract text → AI analyzes → structured data.
  • Classification, tagging, summarization, metadata extraction, all automated.
  • Works for invoices, contracts, reports, and correspondence.
  • Local with Ollama: all documents stay private.
  • Watch out for OCR errors and prompt injection.

FAQ

What is document analysis with AI?

The AI reads a document and understands its content: classification, tagging, summarization, metadata extraction. Unstructured text becomes structured data.

What file formats are supported?

PDF (with text layer), Word, plain text. For scanned PDFs or images, run OCR first (Tesseract or Vision model).

How accurate is the analysis?

Very good for standard documents. For unusual documents or poor OCR quality, the AI can be wrong. For critical data, use human review.

Are my documents secure?

Yes, if you use Ollama locally. All documents stay on your machine. With cloud APIs, documents leave your system. For confidential documents, use local AI.

What does it cost?

Free. Ollama is open source. You only pay for hardware. No per-document API costs.

Can I analyze many documents at once?

Yes, with a pipeline: read documents → extract text → AI analyzes → save results. For many documents, process asynchronously.

Sources and further reading

Back to Blog
Share:

Related Posts