Skip to content
BotServBotServ
OCRTesseractVision modelsText extractionScans

OCR with Local AI

Extract text from images and scans using local AI. Compare Tesseract vs. vision models with practical examples.

S

schutzgeist

4 min read
OCR with Local AI

OCR with Local AI

What this article covers

  • How OCR works and why AI helps.
  • Tesseract (classical) vs. Vision models (modern).
  • How to convert scanned documents, photos, and screenshots into text.
  • Practical examples for document processing.
  • Best practices for accuracy and performance.

Introduction: OCR with local AI explained

OCR (Optical Character Recognition) extracts text from images. Classical OCR (Tesseract) recognizes characters but lacks context. AI-based OCR (Vision models) understands the image: “This is a table with prices” instead of just “A, B, 1, 2”. Keeping everything local with Ollama ensures your images stay private.

This article is for anyone who needs to extract text from images. You’ll find background information in Vision Models and Document Analysis.

Why do you need OCR with AI?

Imagine you have a photo of a document. Tesseract extracts text but misses context: “RE 2024 1234 15 03 1234 56”. A Vision model understands: “Invoice RE-2024-1234, date 15.03., amount 1,234.56 €”. The AI interprets the image, not just the characters.

OCR with local AI at a glance

Image → Vision Model (Ollama) → Text + Context. Or: Image → Tesseract → Raw text → LLM → structured data. Both approaches run locally.

The key insight: classical OCR extracts characters, AI understands content.

Who should read this?

  • Document processors digitizing scans.
  • Developers building OCR pipelines.
  • Archivists capturing legacy documents.
  • Self-hosters running OCR locally.

Key terms

  • OCR - Optical Character Recognition. Useful for: extracting text from images.
  • Tesseract - Classical OCR engine. Useful for: simple extraction.
  • Vision Models - AI for images. Useful for: contextual OCR.
  • Ollama - Model server. Useful for: Vision models.
  • PDF/Text layer - Text embedded in PDFs. Useful for: direct extraction.

Tesseract vs. Vision Model

AspectTesseractVision Model
WhatCharacter recognitionImage understanding
ContextNoYes
TablesPoorGood
HandwritingLimitedBetter
LayoutLoses structurePreserves structure
SpeedVery fastSlower
VRAMNone2-6 GB
Best forSimple textComplex documents

Setup: Tesseract

# Install Tesseract
sudo apt install tesseract-ocr tesseract-ocr-deu

# Python library
pip install pytesseract pillow

# Test
tesseract --version
from PIL import Image
import pytesseract

# Extract text from image
image = Image.open("scan.png")
text = pytesseract.image_to_string(image, lang="deu")
print(text)

Setup: Vision Model (Ollama)

# Pull Vision model
ollama pull minicpm-v:8b
import requests
import base64

def ocr_with_vision(image_path):
    """OCR with Vision model"""
    with open(image_path, "rb") as f:
        image_b64 = base64.b64encode(f.read()).decode()

    response = requests.post("http://ollama:11434/api/chat", json={
        "model": "minicpm-v:8b",
        "messages": [
            {
                "role": "user",
                "content": "Extract all text from this image. Return the text in a structured format.",
                "images": [image_b64]
            }
        ],
        "stream": False
    })
    return response.json()["message"]["content"]

# Example
text = ocr_with_vision("scan.png")

Practical example 1: Scanning an invoice

# Tesseract for quick text
raw_text = pytesseract.image_to_string(invoice_image, lang="deu")

# LLM for structuring
def structure_invoice(raw_text):
    response = requests.post("http://ollama:11434/api/chat", json={
        "model": "llama3.1",
        "messages": [
            {"role": "system", "content": "Structure the OCR text as JSON: invoice_number, date, amount, items."},
            {"role": "user", "content": f"OCR text:\n{raw_text}"}
        ],
        "stream": False,
        "format": "json"
    })
    return json.loads(response.json()["message"]["content"])

Practical example 2: Extracting tables

# Vision model for tables
def extract_table(image_path):
    response = requests.post("http://ollama:11434/api/chat", json={
        "model": "minicpm-v:8b",
        "messages": [
            {
                "role": "user",
                "content": "Extract the table from this image. Return it as a Markdown table.",
                "images": [image_b64]
            }
        ],
        "stream": False
    })
    return response.json()["message"]["content"]

Practical example 3: Handwriting

# Vision model for handwriting (better than Tesseract)
def extract_handwriting(image_path):
    response = requests.post("http://ollama:11434/api/chat", json={
        "model": "minicpm-v:8b",
        "messages": [
            {
                "role": "user",
                "content": "Read the handwriting in this image. Return the text.",
                "images": [image_b64]
            }
        ],
        "stream": False
    })
    return response.json()["message"]["content"]

Hybrid approach

def hybrid_ocr(image_path):
    """Tesseract + LLM for best results"""
    # 1. Tesseract for quick text
    raw_text = pytesseract.image_to_string(image_path, lang="deu")

    # 2. LLM for structuring and error correction
    response = requests.post("http://ollama:11434/api/chat", json={
        "model": "llama3.1",
        "messages": [
            {"role": "system", "content": "Correct OCR errors and structure the text."},
            {"role": "user", "content": f"OCR raw text:\n{raw_text}"}
        ],
        "stream": False
    })
    return response.json()["message"]["content"]

Security notes

  • Images can contain sensitive data: local processing keeps them safe. See Data Protection.
  • OCR errors: poor scans lead to poor extraction. Check quality first.
  • Prompt Injection: images can contain hidden instructions. See Prompt Injection.

Common pitfalls

  • Poor image quality: blurry scans produce bad OCR. Improve quality first.
  • Wrong tool: Tesseract is faster for simple text. Use Vision models for complex documents.
  • No context: Tesseract doesn’t understand context. Chain an LLM for structure.
  • Multiple languages: Tesseract needs language packs. Vision models handle multiple languages automatically.
  • Handwriting: Tesseract struggles with handwriting. Vision models perform better.

Further reading

Key takeaways:

  • Tesseract for fast, simple text extraction.
  • Vision models for complex documents, tables, handwriting.
  • Hybrid: Tesseract for raw text, LLM for structuring.
  • Local with Ollama: all images stay private.
  • For poor scans: improve image quality first.

FAQ

What is OCR?

Optical Character Recognition: extracting text from images. Classically with Tesseract, modernly with Vision models that also understand context.

Tesseract or Vision model?

Tesseract for fast, simple text. Vision model for complex documents, tables, handwriting, and context understanding. Hybrid for best results.

Can AI read handwriting?

Yes, Vision models can read handwriting better than Tesseract. Very poor handwriting is still challenging.

Can AI extract tables?

Yes, Vision models understand table structure and can return them as Markdown or JSON. Tesseract loses structure.

Which languages are supported?

Tesseract requires language packs (tesseract-ocr-deu for German). Vision models are multilingual and handle many languages automatically.

How important is image quality?

Very important. Blurry, skewed, or dark scans produce poor results. For best OCR: high resolution, straight alignment, good contrast.

Are my images safe?

Yes, if you use local tools (Tesseract, Ollama). With cloud OCR (Google, Azure), images leave your system. For sensitive documents, keep everything local.

What does OCR cost?

Tesseract: free. Ollama + Vision model: free (hardware only). Cloud OCR: charged per page or image.

Sources and further reading

Back to Blog
Share:

Related Posts