Document Analysis with Local AI
What this article covers
- How to analyze documents using local AI.
- How text extraction, summarization, and metadata extraction work.
- How to classify and tag documents automatically.
- Practical examples for invoices, contracts, reports, and correspondence.
- Best practices for accuracy, privacy, and performance.
Introduction: Document analysis explained
Document analysis means the AI reads a document and understands its content. It can summarize, classify, extract metadata, and answer questions. Using Ollama locally keeps all documents private.
This article is for anyone who wants to analyze documents with AI. You’ll find foundational material in Documents and PDF and Ollama.
Why do you need document analysis?
Imagine you have 500 PDFs: invoices, contracts, reports. Instead of reading each one manually, the AI analyzes: “This is an invoice from Company X for 1,234 EUR, due on March 15. Risk: no payment terms specified.” The AI transforms unstructured documents into structured data.
Document analysis at a glance
Document → Extract text (OCR if needed) → AI analyzes → Structured data (category, tags, summary, metadata). All handled locally by Ollama.
The core concept: turn unstructured text into structured data.
Who is this article for?
- Office workers who need to process documents automatically.
- Businesses analyzing invoices and contracts.
- Developers building document pipelines.
- Self-hosters running document analysis locally.
Key terms
- Ollama - Local model server. Useful for: the AI backend.
- OCR - Text from images. Useful for: scanned documents.
- Classification - Categorizing documents. Useful for: organizing files.
- Metadata - Structured data. Useful for: databases.
- RAG - Document Q&A. Useful for: asking questions.
- Vision models - For images. Useful for: scanning documents.
The pipeline
Document (PDF/Word/Image)
│
▼
Text Extraction
├─ Text present? → Use directly
└─ No text? → OCR (Tesseract/Vision model)
│
▼
AI Analysis (Ollama)
├─ Classify: Invoice? Contract? Report?
├─ Generate tags: topics, categories
├─ Summarize: key content
└─ Extract metadata: date, sender, amount, ...
│
▼
Structured Data
├─ Store in database
├─ Move to folder
└─ Export as JSON
Practical example 1: Analyzing invoices
import requests
import json
def analyze_invoice(text):
"""Analyze invoice"""
response = requests.post("http://ollama:11434/api/chat", json={
"model": "llama3.1",
"messages": [
{"role": "system", "content": "Extract invoice data as JSON: invoice_number, date, amount, currency, sender, due_date, line_items (array)."},
{"role": "user", "content": f"Invoice:\n{text[:4000]}"}
],
"stream": False,
"format": "json"
})
return json.loads(response.json()["message"]["content"])
# Example
invoice_data = analyze_invoice(invoice_text)
# Output: {"invoice_number": "RE-2024-1234", "amount": 1234.56, "currency": "EUR", ...}
Practical example 2: Analyzing contracts
def analyze_contract(text):
"""Analyze contract"""
response = requests.post("http://ollama:11434/api/chat", json={
"model": "llama3.1",
"messages": [
{"role": "system", "content": "Analyze the contract as JSON: contract_type, parties, duration, termination_notice, clauses (array), risks (array)."},
{"role": "user", "content": f"Contract:\n{text[:6000]}"}
],
"stream": False,
"format": "json"
})
return json.loads(response.json()["message"]["content"])
Practical example 3: Classifying documents
def classify_document(text):
"""Classify document"""
categories = [
"invoice", "contract", "report", "correspondence",
"minutes", "manual", "other"
]
response = requests.post("http://ollama:11434/api/chat", json={
"model": "llama3.1",
"messages": [
{"role": "system", "content": f"Classify the document into one of these categories: {', '.join(categories)}. Respond with only the category name."},
{"role": "user", "content": f"Document:\n{text[:2000]}"}
],
"stream": False
})
return response.json()["message"]["content"].strip().lower()
Practical example 4: Generating tags
def generate_tags(text, max_tags=5):
"""Generate tags for document"""
response = requests.post("http://ollama:11434/api/chat", json={
"model": "llama3.1",
"messages": [
{"role": "system", "content": f"Assign {max_tags} relevant tags for the document. Respond as a JSON array."},
{"role": "user", "content": f"Document:\n{text[:2000]}"}
],
"stream": False,
"format": "json"
})
return json.loads(response.json()["message"]["content"])
Practical example 5: Complete pipeline
import os
from pathlib import Path
def process_document(filepath):
"""Complete document analysis"""
# 1. Extract text
text = extract_text(filepath) # PDF, DOCX, TXT
# 2. Classify
category = classify_document(text)
# 3. Generate tags
tags = generate_tags(text)
# 4. Summarize
summary = summarize(text)
# 5. Extract metadata
metadata = extract_metadata(text)
# 6. Store results
result = {
"file": filepath,
"category": category,
"tags": tags,
"summary": summary,
"metadata": metadata
}
save_to_database(result)
move_to_folder(filepath, category)
return result
Security considerations
- Confidential documents: All data stays local with Ollama. Cloud APIs send documents outbound. See Privacy.
- Prompt injection: Documents can contain injections. See Prompt Injection.
- OCR errors: Scanned documents have OCR errors. AI results can be incorrect as a result.
- Validation: AI extractions should be validated for critical data.
Common pitfalls
- OCR errors: Poor scans mean poor extraction. Quality matters.
- Context length: Large documents need to be chunked. See Context length.
- Misclassification: AI can classify incorrectly. For critical documents, human review is needed.
- Too many tags: More than 5-7 tags become unwieldy. Set a limit.
- Performance: Large documents plus multiple AI calls slow things down. Process asynchronously.
Further reading
- Documents and PDF - Overview.
- PDF chatbot - Chat with documents.
- Local RAG - Document Q&A.
- Ollama - Model server.
- Prompt Injection - Security.
- Privacy - Data protection.
Key takeaways:
- Document analysis: extract text → AI analyzes → structured data.
- Classification, tagging, summarization, metadata extraction, all automated.
- Works for invoices, contracts, reports, and correspondence.
- Local with Ollama: all documents stay private.
- Watch out for OCR errors and prompt injection.
FAQ
What is document analysis with AI?
What file formats are supported?
How accurate is the analysis?
Are my documents secure?
What does it cost?
Can I analyze many documents at once?
Sources and further reading
- Ollama - Local model server.
- Tesseract OCR - OCR engine.
- PyPDF2 - PDF text extraction.
- python-docx - Word documents.


