Skip to content
BotServBotServ
RAGPDFOCRTesseracttext extractionlocal AIdocuments

PDF and OCR: Scanned Documents for RAG

Extract text from PDFs and scanned documents for RAG using OCR, Tesseract, layout detection, and best practices.

S

schutzgeist

12 min read
PDF and OCR: Scanned Documents for RAG

PDFs and OCR: Processing Scanned Documents for RAG

What This Article Covers

  • The different types of PDFs and when OCR is necessary
  • How to extract text from digital PDFs using PyPDF2, pdfplumber, and PyMuPDF
  • Using Tesseract and pytesseract to convert scanned documents into searchable text
  • Why layout detection and preprocessing dramatically improve OCR quality
  • A complete practical example: converting a scanned manual from image to searchable text

Introduction: Understanding PDFs and OCR

You’re building a local RAG system and have collected your documents. There are manuals in PDF, old contracts as scans, notes as photos. You try extracting the text and realize that many files yield nothing useful. The PDF contains no text, only images.

This is where PDF processing and OCR come in. In this article, I’ll show you how to prepare PDFs and scanned documents so they work in your RAG pipeline. If you haven’t yet covered the document preparation basics, start there. Here we dive deeper into PDF-specific challenges.

Why Do You Need PDF and OCR Processing?

Imagine a folder full of old contracts, user manuals, and technical documentation. Most of them are scanned PDFs, essentially photos of paper pages packaged in a PDF file. There’s no text layer, nothing is searchable, nothing can be copied.

You try loading these documents into your RAG system. Extraction returns empty pages or throws errors. Your RAG system finds no answers because there’s simply no text available. The documents exist, but they’re invisible to your AI.

Without OCR, all those scanned documents remain unused. You’re sitting on a mountain of information but can’t access it. With OCR, you transform those images into searchable text and make them available to RAG. This is the critical step when your document collection includes scanned materials.

A Quick Explanation of PDFs and OCR

Think of it this way: imagine a photocopier that doesn’t just copy, but can also read what it’s copying. A regular copier produces an image of the page. OCR is like a copier that takes that image and simultaneously recognizes the text on it, as if someone were typing out the page.

With digital PDFs, this isn’t necessary. Text is stored as text, like in a word processor document. You can extract it directly. With scanned PDFs, each page is just an image. OCR must recognize the text from that image, character by character, word by word.

Who This Article Is For

This article is for beginners building a local RAG system and working with scanned documents or PDFs. You don’t need prior OCR experience, but should understand the basics of what local AI is and how RAG works. The code examples are in Python and explained so you can follow along without deep programming knowledge.

Key Concepts in PDF and OCR

TermExplanation
PDFA document format that can contain text and images. Not every PDF contains selectable text.
Text LayerThe text information embedded in a PDF. Present in digital PDFs, absent in scanned ones.
OCROptical Character Recognition. Technology that recognizes text from images.
TesseractAn open-source OCR engine that runs locally and supports many languages.
LayoutThe arrangement of text, images, and tables on a page. Critical for reading order.
DPIDots per Inch. The resolution of an image. Higher DPI means more detail for OCR.
Image FormatThe image format within a PDF, such as JPEG or PNG. Affects quality and file size.
PageA single page of a PDF. Processed individually during extraction.
ColumnA vertical text area in a multi-column layout. Breaks reading order if ignored.
ParagraphA cohesive text block. Later becomes a chunk for RAG.
TokenA text fragment processed by models. Relevant for chunking and embeddings.

Types of PDFs

Before you start, you need to understand what type of PDF you’re dealing with. Processing differs fundamentally.

Text-based PDFs (born digital)

These PDFs were exported from a word processor or software. Text is stored as text, not as an image. You can select it, copy it, and search it. Extraction is straightforward and requires no OCR. Examples: PDFs from Word, LaTeX, HTML exports.

Scanned PDFs (image only)

These PDFs consist of images of paper pages. Each page is a photo with no text layer. You can’t select or search the text. You need OCR to extract it. Examples: scanned contracts, old manuals, archived documents.

Mixed PDFs

Some PDFs contain both text and scanned pages. For instance, a manual where the first pages were created digitally but old appendices are included as scans. You need to check each page for text and apply OCR only where necessary.

Here’s how to check in Python whether a page contains text:

from pypdf import PdfReader

reader = PdfReader("dokument.pdf")
for i, page in enumerate(reader.pages):
    text = page.extract_text()
    if text and text.strip():
        print(f"Seite {i+1}: Text vorhanden ({len(text)} Zeichen)")
    else:
        print(f"Seite {i+1}: Kein Text, OCR nötig")

With this check upfront, you save OCR processing on pages that don’t need it.

Extracting Text-Based PDFs

For text-based PDFs, several good Python libraries exist. Here are the three main ones with their strengths.

PyPDF2 (pypdf)

PyPDF2 is the simplest option. It extracts plain text but ignores layout information.

from pypdf import PdfReader

reader = PdfReader("handbuch.pdf")
text = ""
for page in reader.pages:
    text += page.extract_text() + "\n"
print(text[:500])

PyPDF2 is fast and reliable for straightforward PDFs. With multi-column layouts, reading order can become jumbled because layout isn’t considered.

pdfplumber

pdfplumber is better suited for complex layouts. It recognizes columns and tables, giving you more control over extraction.

import pdfplumber

with pdfplumber.open("handbuch.pdf") as pdf:
    for page in pdf.pages:
        text = page.extract_text()
        if text:
            print(text[:500])

pdfplumber also offers methods like extract_tables() to read tables in a structured way. This is valuable if your PDFs contain tabular data.

PyMuPDF (fitz)

PyMuPDF is the fastest option and extracts text with layout information.

import fitz  # PyMuPDF

doc = fitz.open("handbuch.pdf")
for page in doc:
    text = page.get_text()
    print(text[:500])

PyMuPDF also gives you metadata like font size and position, which helps distinguish headings from body text in complex layouts.

Which library you use depends on your documents. For simple PDFs, PyPDF2 suffices. For complex layouts, use pdfplumber or PyMuPDF.

Scanned PDFs with OCR

If a PDF lacks a text layer, you’ll need OCR. Tesseract is the best open-source option for local processing.

Installing Tesseract

On Ubuntu or Debian:

sudo apt install tesseract-ocr tesseract-ocr-deu

On macOS with Homebrew:

brew install tesseract tesseract-lang

On Windows, download the installer from github.com/UB-Mannheim/tesseract.

The tesseract-ocr-deu package installs the German language model. English uses tesseract-ocr-eng. You can install multiple languages simultaneously.

Installing pytesseract

pip install pytesseract Pillow pdf2image

pytesseract is the Python interface for Tesseract. Pillow handles image processing, and pdf2image converts PDF pages into images.

Converting PDFs to images and running OCR

from pdf2image import convert_from_path
import pytesseract

# Convert PDF to images at 300 DPI for good OCR quality
images = convert_from_path("scan.pdf", dpi=300)

full_text = ""
for i, image in enumerate(images):
    # OCR with German language model
    text = pytesseract.image_to_string(image, lang="deu")
    full_text += f"--- Page {i+1} ---\n{text}\n"

print(full_text[:500])

Setting lang="deu" tells Tesseract the text is in German. For English documents, use lang="eng". You can also specify multiple languages: lang="deu+eng".

Layout Recognition

Basic OCR recognizes characters but understands nothing about page structure. With multi-column layouts, tables, or embedded images, the reading order becomes scrambled. The text becomes unusable for RAG because related sentences get split apart.

Why layout matters

Imagine a newspaper page with three columns. Basic OCR reads line by line across all columns at once, producing gibberish. The first column must be read completely before moving to the second. Only then does the meaning survive.

LayoutLM

LayoutLM is a Microsoft model that processes text and layout information together. It recognizes which text blocks belong together and the order they should be read. LayoutLM is powerful but requires more setup.

PaddleOCR

PaddleOCR is another open-source option combining layout recognition and OCR. It’s simpler to set up than LayoutLM and delivers good results on multi-column layouts and tables.

from paddleocr import PaddleOCR

ocr = PaddleOCR(use_angle_cls=True, lang="german")
result = ocr.ocr("scan.png", cls=True)

for line in result[0]:
    print(line[1][0])

Tesseract works fine for simple documents. If your documents have complex layouts, PaddleOCR or LayoutLM are worth the effort.

Improving Quality

OCR quality depends heavily on input image quality. Preprocessing significantly improves recognition rates.

Increasing DPI

Scan at least 300 DPI, preferably 600 DPI for small text. Lower resolution causes errors because Tesseract can’t cleanly recognize characters.

Binarization

Binarization converts images to black and white, making Tesseract’s job easier.

from PIL import Image, ImageOps

image = Image.open("scan.png")
# Convert to grayscale
gray = ImageOps.grayscale(image)
# Binarize with threshold
bw = gray.point(lambda x: 0 if x < 128 else 255, "1")
bw.save("scan_bw.png")

Deskewing

Skewed scans degrade OCR quality. Use deskew to straighten the image.

from skimage import io
from skimage.transform import rotate
from skimage.color import rgb2gray
import numpy as np
from scipy.ndimage import interpolation

def deskew(image_path):
    image = io.imread(image_path, as_gray=True)
    # Threshold for binarization
    binary = image < 0.5
    # Estimate angle
    angles = np.linspace(-10, 10, 41)
    scores = []
    for angle in angles:
        rotated = interpolation.rotate(binary, angle, reshape=False)
        scores.append(np.sum(rotated))
    best_angle = angles[np.argmax(scores)]
    # Straighten image
    corrected = rotate(image, best_angle, resize=True)
    io.imsave("scan_deskewed.png", corrected)

deskew("scan.png")

These three steps, higher DPI, binarization, and deskewing, can drastically reduce OCR error rates. Investing time in preprocessing pays off in every subsequent RAG query.

Example: Processing a Scanned Manual

Let’s work through a complete example. You have a scanned manual as a PDF and want to prepare it for RAG.

Step 1: Convert PDF to images

from pdf2image import convert_from_path

images = convert_from_path("handbuch_scan.pdf", dpi=300)
print(f"{len(images)} pages loaded")

Step 2: Apply preprocessing

from PIL import Image, ImageOps, ImageFilter

def preprocess(image):
    # Grayscale
    gray = ImageOps.grayscale(image)
    # Increase contrast
    gray = ImageOps.autocontrast(gray)
    # Light sharpening
    gray = gray.filter(ImageFilter.SHARPEN)
    # Binarization
    bw = gray.point(lambda x: 0 if x < 140 else 255, "1")
    return bw

processed_images = [preprocess(img) for img in images]

Step 3: Run OCR

import pytesseract

pages = []
for i, img in enumerate(processed_images):
    text = pytesseract.image_to_string(img, lang="deu")
    pages.append({"text": text, "page": i + 1})
    print(f"Page {i+1}: {len(text)} characters recognized")

Step 4: Clean the text

import re

def clean_ocr_text(text):
    # Normalize whitespace
    text = re.sub(r"[ \t]+", " ", text)
    # Reduce blank lines
    text = re.sub(r"\n{3,}", "\n\n", text)
    # Fix common OCR errors
    text = text.replace("|", "l")
    text = text.replace("0", "o")  # Use with caution!
    return text.strip()

for page in pages:
    page["text"] = clean_ocr_text(page["text"])

Be careful with automatic error correction: replacements like 0 to o can change correct characters too. Inspect results and tailor corrections to your documents.

Step 5: Add metadata and create chunks

chunks = []
for page in pages:
    paragraphs = page["text"].split("\n\n")
    for para in paragraphs:
        if len(para.strip()) > 50:
            chunks.append({
                "text": para.strip(),
                "metadata": {
                    "source": "handbuch_scan.pdf",
                    "page": page["page"],
                    "method": "ocr"
                }
            })

print(f"{len(chunks)} chunks created")

From here, continue as described in the document preparation article: generate embeddings and load the chunks into your vector database.

OCR Alternatives

Tesseract isn’t your only option. Here are alternatives, especially if you need higher quality.

Ollama with Vision Models

With Ollama, you can run vision models like LLaVA locally. These models can “read” images and extract text from them. They’re more flexible than Tesseract, especially for handwritten notes or unusual layouts, but slower.

# Pseudocode, see Ollama documentation for details
import ollama

response = ollama.chat(
    model="llava",
    messages=[{
        "role": "user",
        "content": "Extract the text from this image.",
        "images": ["scan.png"]
    }]
)
print(response["message"]["content"])

Vision models complement Tesseract well when traditional OCR hits its limits.

Cloud OCR

Services like Google Cloud Vision or AWS Textract offer excellent OCR quality. The downside: your documents leave your machine. For sensitive documents, that’s often unacceptable. If you’re reading this article, you’re probably interested in local AI, so we’ll stick with local solutions.

Common pitfalls with PDFs and OCR

  1. Resolution too low: Scans below 200 DPI produce severe OCR errors. Use at least 300 DPI.
  2. Wrong language model: Running Tesseract with an English model on German text yields unusable results. Check your language settings.
  3. Layout not recognized: Multi-column documents processed without layout detection become unreadable. Use PaddleOCR or LayoutLM for complex layouts.
  4. No preprocessing: Raw scans without binarization and deskewing degrade quality significantly. Invest time in preprocessing.
  5. OCR errors not corrected: OCR is never perfect. Common mistakes like rn instead of m or 1 instead of l hurt RAG quality. Review and fix them.
  6. Page numbers and headers not removed: Clean up headers and footers from OCR text, as described in the document preparation article.
  7. Images too large in memory: 600-DPI scans of an A4 page consume significant RAM. Process pages individually, not all at once.
  8. Mixed PDFs not detected: Applying OCR to text-based pages degrades results. Check whether text already exists first.

Hardware, costs, and security for PDFs and OCR

OCR is computationally intensive, but not extreme. Tesseract runs smoothly on a standard laptop without a GPU. Processing 100 pages takes roughly 5 to 15 minutes on an average CPU, depending on DPI and complexity.

Costs are essentially zero. Tesseract is open source and runs locally. You need no API keys and no cloud subscriptions. Storage for processed images is the only resource factor.

Security is a major advantage of local OCR. Your documents never leave your machine. This matters especially for contracts, medical records, and other sensitive files. You avoid data leakage risks that come with cloud OCR services.

Further reading and resources on PDFs and OCR

FAQ: PDFs and OCR - Common questions

Do I need OCR for every PDF?

No. Only for scanned PDFs without a text layer. Digital PDFs exported from word processors contain selectable text. Check first with PyPDF2 whether text is extractable.

What’s better: Tesseract or a vision model?

For standard printed text, Tesseract is faster and sufficient. For handwritten notes, unusual layouts, or text with many special characters, a vision model like LLaVA over Ollama may yield better results, though it’s slower.

What DPI is optimal for OCR?

300 DPI is a good default. For small text or fine details, use 600 DPI. Below 200 DPI, the error rate rises noticeably.

How do I handle multi-column layouts?

Use a tool with layout recognition like PaddleOCR or LayoutLM. Basic Tesseract reads line by line across all columns and produces unusable text.

Can Tesseract recognize handwriting?

Tesseract is optimized for printed text. It detects handwriting unreliably. For handwritten notes, vision models work better.

How do I correct OCR errors?

Compare OCR results to the original, identify common error patterns, and apply targeted replacements. For large volumes, you can use a language model for post-correction that spots obvious errors in context.

How long does OCR take for 100 pages?

On an average CPU, roughly 5 to 15 minutes at 300 DPI. With a GPU it’s much faster. Duration also depends on page complexity.

Do I need a GPU for OCR?

No, Tesseract runs on the CPU. For vision models via Ollama, a GPU is recommended but not required. Processing is simply slower without one.

What about tables in scanned documents?

Tables challenge OCR engines. PaddleOCR and LayoutLM can detect them. Alternatively, keep tables as images and extract them with a vision model. For RAG, storing tables as structured metadata rather than flowing text often works better.

Can I process multiple languages at once?

Yes, Tesseract supports multiple languages simultaneously with lang="deu+eng". This helps with documents mixing German and English text. Recognition accuracy may drop slightly compared to single-language processing.

Sources and further reading

  • Tesseract documentation: github.com/tesseract-ocr/tesseract
  • pytesseract documentation: github.com/madmaze/pytesseract
  • pdfplumber documentation: github.com/jsvine/pdfplumber
  • PyMuPDF documentation: pymupdf.readthedocs.io
  • PaddleOCR: github.com/PaddlePaddle/PaddleOCR
  • LayoutLM: github.com/microsoft/unilm
  • Article on document preparation in this series
  • Article on RAG fundamentals in this series
Back to Blog
Share:

Related Posts