PDFs and OCR: Processing Scanned Documents for RAG
What This Article Covers
- The different types of PDFs and when OCR is necessary
- How to extract text from digital PDFs using PyPDF2, pdfplumber, and PyMuPDF
- Using Tesseract and pytesseract to convert scanned documents into searchable text
- Why layout detection and preprocessing dramatically improve OCR quality
- A complete practical example: converting a scanned manual from image to searchable text
Introduction: Understanding PDFs and OCR
You’re building a local RAG system and have collected your documents. There are manuals in PDF, old contracts as scans, notes as photos. You try extracting the text and realize that many files yield nothing useful. The PDF contains no text, only images.
This is where PDF processing and OCR come in. In this article, I’ll show you how to prepare PDFs and scanned documents so they work in your RAG pipeline. If you haven’t yet covered the document preparation basics, start there. Here we dive deeper into PDF-specific challenges.
Why Do You Need PDF and OCR Processing?
Imagine a folder full of old contracts, user manuals, and technical documentation. Most of them are scanned PDFs, essentially photos of paper pages packaged in a PDF file. There’s no text layer, nothing is searchable, nothing can be copied.
You try loading these documents into your RAG system. Extraction returns empty pages or throws errors. Your RAG system finds no answers because there’s simply no text available. The documents exist, but they’re invisible to your AI.
Without OCR, all those scanned documents remain unused. You’re sitting on a mountain of information but can’t access it. With OCR, you transform those images into searchable text and make them available to RAG. This is the critical step when your document collection includes scanned materials.
A Quick Explanation of PDFs and OCR
Think of it this way: imagine a photocopier that doesn’t just copy, but can also read what it’s copying. A regular copier produces an image of the page. OCR is like a copier that takes that image and simultaneously recognizes the text on it, as if someone were typing out the page.
With digital PDFs, this isn’t necessary. Text is stored as text, like in a word processor document. You can extract it directly. With scanned PDFs, each page is just an image. OCR must recognize the text from that image, character by character, word by word.
Who This Article Is For
This article is for beginners building a local RAG system and working with scanned documents or PDFs. You don’t need prior OCR experience, but should understand the basics of what local AI is and how RAG works. The code examples are in Python and explained so you can follow along without deep programming knowledge.
Key Concepts in PDF and OCR
| Term | Explanation |
|---|---|
| A document format that can contain text and images. Not every PDF contains selectable text. | |
| Text Layer | The text information embedded in a PDF. Present in digital PDFs, absent in scanned ones. |
| OCR | Optical Character Recognition. Technology that recognizes text from images. |
| Tesseract | An open-source OCR engine that runs locally and supports many languages. |
| Layout | The arrangement of text, images, and tables on a page. Critical for reading order. |
| DPI | Dots per Inch. The resolution of an image. Higher DPI means more detail for OCR. |
| Image Format | The image format within a PDF, such as JPEG or PNG. Affects quality and file size. |
| Page | A single page of a PDF. Processed individually during extraction. |
| Column | A vertical text area in a multi-column layout. Breaks reading order if ignored. |
| Paragraph | A cohesive text block. Later becomes a chunk for RAG. |
| Token | A text fragment processed by models. Relevant for chunking and embeddings. |
Types of PDFs
Before you start, you need to understand what type of PDF you’re dealing with. Processing differs fundamentally.
Text-based PDFs (born digital)
These PDFs were exported from a word processor or software. Text is stored as text, not as an image. You can select it, copy it, and search it. Extraction is straightforward and requires no OCR. Examples: PDFs from Word, LaTeX, HTML exports.
Scanned PDFs (image only)
These PDFs consist of images of paper pages. Each page is a photo with no text layer. You can’t select or search the text. You need OCR to extract it. Examples: scanned contracts, old manuals, archived documents.
Mixed PDFs
Some PDFs contain both text and scanned pages. For instance, a manual where the first pages were created digitally but old appendices are included as scans. You need to check each page for text and apply OCR only where necessary.
Here’s how to check in Python whether a page contains text:
from pypdf import PdfReader
reader = PdfReader("dokument.pdf")
for i, page in enumerate(reader.pages):
text = page.extract_text()
if text and text.strip():
print(f"Seite {i+1}: Text vorhanden ({len(text)} Zeichen)")
else:
print(f"Seite {i+1}: Kein Text, OCR nötig")
With this check upfront, you save OCR processing on pages that don’t need it.
Extracting Text-Based PDFs
For text-based PDFs, several good Python libraries exist. Here are the three main ones with their strengths.
PyPDF2 (pypdf)
PyPDF2 is the simplest option. It extracts plain text but ignores layout information.
from pypdf import PdfReader
reader = PdfReader("handbuch.pdf")
text = ""
for page in reader.pages:
text += page.extract_text() + "\n"
print(text[:500])
PyPDF2 is fast and reliable for straightforward PDFs. With multi-column layouts, reading order can become jumbled because layout isn’t considered.
pdfplumber
pdfplumber is better suited for complex layouts. It recognizes columns and tables, giving you more control over extraction.
import pdfplumber
with pdfplumber.open("handbuch.pdf") as pdf:
for page in pdf.pages:
text = page.extract_text()
if text:
print(text[:500])
pdfplumber also offers methods like extract_tables() to read tables in a structured way. This is valuable if your PDFs contain tabular data.
PyMuPDF (fitz)
PyMuPDF is the fastest option and extracts text with layout information.
import fitz # PyMuPDF
doc = fitz.open("handbuch.pdf")
for page in doc:
text = page.get_text()
print(text[:500])
PyMuPDF also gives you metadata like font size and position, which helps distinguish headings from body text in complex layouts.
Which library you use depends on your documents. For simple PDFs, PyPDF2 suffices. For complex layouts, use pdfplumber or PyMuPDF.
Scanned PDFs with OCR
If a PDF lacks a text layer, you’ll need OCR. Tesseract is the best open-source option for local processing.
Installing Tesseract
On Ubuntu or Debian:
sudo apt install tesseract-ocr tesseract-ocr-deu
On macOS with Homebrew:
brew install tesseract tesseract-lang
On Windows, download the installer from github.com/UB-Mannheim/tesseract.
The tesseract-ocr-deu package installs the German language model. English uses tesseract-ocr-eng. You can install multiple languages simultaneously.
Installing pytesseract
pip install pytesseract Pillow pdf2image
pytesseract is the Python interface for Tesseract. Pillow handles image processing, and pdf2image converts PDF pages into images.
Converting PDFs to images and running OCR
from pdf2image import convert_from_path
import pytesseract
# Convert PDF to images at 300 DPI for good OCR quality
images = convert_from_path("scan.pdf", dpi=300)
full_text = ""
for i, image in enumerate(images):
# OCR with German language model
text = pytesseract.image_to_string(image, lang="deu")
full_text += f"--- Page {i+1} ---\n{text}\n"
print(full_text[:500])
Setting lang="deu" tells Tesseract the text is in German. For English documents, use lang="eng". You can also specify multiple languages: lang="deu+eng".
Layout Recognition
Basic OCR recognizes characters but understands nothing about page structure. With multi-column layouts, tables, or embedded images, the reading order becomes scrambled. The text becomes unusable for RAG because related sentences get split apart.
Why layout matters
Imagine a newspaper page with three columns. Basic OCR reads line by line across all columns at once, producing gibberish. The first column must be read completely before moving to the second. Only then does the meaning survive.
LayoutLM
LayoutLM is a Microsoft model that processes text and layout information together. It recognizes which text blocks belong together and the order they should be read. LayoutLM is powerful but requires more setup.
PaddleOCR
PaddleOCR is another open-source option combining layout recognition and OCR. It’s simpler to set up than LayoutLM and delivers good results on multi-column layouts and tables.
from paddleocr import PaddleOCR
ocr = PaddleOCR(use_angle_cls=True, lang="german")
result = ocr.ocr("scan.png", cls=True)
for line in result[0]:
print(line[1][0])
Tesseract works fine for simple documents. If your documents have complex layouts, PaddleOCR or LayoutLM are worth the effort.
Improving Quality
OCR quality depends heavily on input image quality. Preprocessing significantly improves recognition rates.
Increasing DPI
Scan at least 300 DPI, preferably 600 DPI for small text. Lower resolution causes errors because Tesseract can’t cleanly recognize characters.
Binarization
Binarization converts images to black and white, making Tesseract’s job easier.
from PIL import Image, ImageOps
image = Image.open("scan.png")
# Convert to grayscale
gray = ImageOps.grayscale(image)
# Binarize with threshold
bw = gray.point(lambda x: 0 if x < 128 else 255, "1")
bw.save("scan_bw.png")
Deskewing
Skewed scans degrade OCR quality. Use deskew to straighten the image.
from skimage import io
from skimage.transform import rotate
from skimage.color import rgb2gray
import numpy as np
from scipy.ndimage import interpolation
def deskew(image_path):
image = io.imread(image_path, as_gray=True)
# Threshold for binarization
binary = image < 0.5
# Estimate angle
angles = np.linspace(-10, 10, 41)
scores = []
for angle in angles:
rotated = interpolation.rotate(binary, angle, reshape=False)
scores.append(np.sum(rotated))
best_angle = angles[np.argmax(scores)]
# Straighten image
corrected = rotate(image, best_angle, resize=True)
io.imsave("scan_deskewed.png", corrected)
deskew("scan.png")
These three steps, higher DPI, binarization, and deskewing, can drastically reduce OCR error rates. Investing time in preprocessing pays off in every subsequent RAG query.
Example: Processing a Scanned Manual
Let’s work through a complete example. You have a scanned manual as a PDF and want to prepare it for RAG.
Step 1: Convert PDF to images
from pdf2image import convert_from_path
images = convert_from_path("handbuch_scan.pdf", dpi=300)
print(f"{len(images)} pages loaded")
Step 2: Apply preprocessing
from PIL import Image, ImageOps, ImageFilter
def preprocess(image):
# Grayscale
gray = ImageOps.grayscale(image)
# Increase contrast
gray = ImageOps.autocontrast(gray)
# Light sharpening
gray = gray.filter(ImageFilter.SHARPEN)
# Binarization
bw = gray.point(lambda x: 0 if x < 140 else 255, "1")
return bw
processed_images = [preprocess(img) for img in images]
Step 3: Run OCR
import pytesseract
pages = []
for i, img in enumerate(processed_images):
text = pytesseract.image_to_string(img, lang="deu")
pages.append({"text": text, "page": i + 1})
print(f"Page {i+1}: {len(text)} characters recognized")
Step 4: Clean the text
import re
def clean_ocr_text(text):
# Normalize whitespace
text = re.sub(r"[ \t]+", " ", text)
# Reduce blank lines
text = re.sub(r"\n{3,}", "\n\n", text)
# Fix common OCR errors
text = text.replace("|", "l")
text = text.replace("0", "o") # Use with caution!
return text.strip()
for page in pages:
page["text"] = clean_ocr_text(page["text"])
Be careful with automatic error correction: replacements like 0 to o can change correct characters too. Inspect results and tailor corrections to your documents.
Step 5: Add metadata and create chunks
chunks = []
for page in pages:
paragraphs = page["text"].split("\n\n")
for para in paragraphs:
if len(para.strip()) > 50:
chunks.append({
"text": para.strip(),
"metadata": {
"source": "handbuch_scan.pdf",
"page": page["page"],
"method": "ocr"
}
})
print(f"{len(chunks)} chunks created")
From here, continue as described in the document preparation article: generate embeddings and load the chunks into your vector database.
OCR Alternatives
Tesseract isn’t your only option. Here are alternatives, especially if you need higher quality.
Ollama with Vision Models
With Ollama, you can run vision models like LLaVA locally. These models can “read” images and extract text from them. They’re more flexible than Tesseract, especially for handwritten notes or unusual layouts, but slower.
# Pseudocode, see Ollama documentation for details
import ollama
response = ollama.chat(
model="llava",
messages=[{
"role": "user",
"content": "Extract the text from this image.",
"images": ["scan.png"]
}]
)
print(response["message"]["content"])
Vision models complement Tesseract well when traditional OCR hits its limits.
Cloud OCR
Services like Google Cloud Vision or AWS Textract offer excellent OCR quality. The downside: your documents leave your machine. For sensitive documents, that’s often unacceptable. If you’re reading this article, you’re probably interested in local AI, so we’ll stick with local solutions.
Common pitfalls with PDFs and OCR
- Resolution too low: Scans below 200 DPI produce severe OCR errors. Use at least 300 DPI.
- Wrong language model: Running Tesseract with an English model on German text yields unusable results. Check your language settings.
- Layout not recognized: Multi-column documents processed without layout detection become unreadable. Use PaddleOCR or LayoutLM for complex layouts.
- No preprocessing: Raw scans without binarization and deskewing degrade quality significantly. Invest time in preprocessing.
- OCR errors not corrected: OCR is never perfect. Common mistakes like
rninstead ofmor1instead oflhurt RAG quality. Review and fix them. - Page numbers and headers not removed: Clean up headers and footers from OCR text, as described in the document preparation article.
- Images too large in memory: 600-DPI scans of an A4 page consume significant RAM. Process pages individually, not all at once.
- Mixed PDFs not detected: Applying OCR to text-based pages degrades results. Check whether text already exists first.
Hardware, costs, and security for PDFs and OCR
OCR is computationally intensive, but not extreme. Tesseract runs smoothly on a standard laptop without a GPU. Processing 100 pages takes roughly 5 to 15 minutes on an average CPU, depending on DPI and complexity.
Costs are essentially zero. Tesseract is open source and runs locally. You need no API keys and no cloud subscriptions. Storage for processed images is the only resource factor.
Security is a major advantage of local OCR. Your documents never leave your machine. This matters especially for contracts, medical records, and other sensitive files. You avoid data leakage risks that come with cloud OCR services.
Further reading and resources on PDFs and OCR
- Local RAG - Overview of all RAG articles
- RAG fundamentals - How RAG works
- Document preparation - Clean data processing
- Chunking - Text splitting done right
- Embedding models - Generate vectors from text
- What is local AI? - Foundations of local AI
- Ollama - Run local models
FAQ: PDFs and OCR - Common questions
Do I need OCR for every PDF?
No. Only for scanned PDFs without a text layer. Digital PDFs exported from word processors contain selectable text. Check first with PyPDF2 whether text is extractable.
What’s better: Tesseract or a vision model?
For standard printed text, Tesseract is faster and sufficient. For handwritten notes, unusual layouts, or text with many special characters, a vision model like LLaVA over Ollama may yield better results, though it’s slower.
What DPI is optimal for OCR?
300 DPI is a good default. For small text or fine details, use 600 DPI. Below 200 DPI, the error rate rises noticeably.
How do I handle multi-column layouts?
Use a tool with layout recognition like PaddleOCR or LayoutLM. Basic Tesseract reads line by line across all columns and produces unusable text.
Can Tesseract recognize handwriting?
Tesseract is optimized for printed text. It detects handwriting unreliably. For handwritten notes, vision models work better.
How do I correct OCR errors?
Compare OCR results to the original, identify common error patterns, and apply targeted replacements. For large volumes, you can use a language model for post-correction that spots obvious errors in context.
How long does OCR take for 100 pages?
On an average CPU, roughly 5 to 15 minutes at 300 DPI. With a GPU it’s much faster. Duration also depends on page complexity.
Do I need a GPU for OCR?
No, Tesseract runs on the CPU. For vision models via Ollama, a GPU is recommended but not required. Processing is simply slower without one.
What about tables in scanned documents?
Tables challenge OCR engines. PaddleOCR and LayoutLM can detect them. Alternatively, keep tables as images and extract them with a vision model. For RAG, storing tables as structured metadata rather than flowing text often works better.
Can I process multiple languages at once?
Yes, Tesseract supports multiple languages simultaneously with lang="deu+eng". This helps with documents mixing German and English text. Recognition accuracy may drop slightly compared to single-language processing.
Sources and further reading
- Tesseract documentation: github.com/tesseract-ocr/tesseract
- pytesseract documentation: github.com/madmaze/pytesseract
- pdfplumber documentation: github.com/jsvine/pdfplumber
- PyMuPDF documentation: pymupdf.readthedocs.io
- PaddleOCR: github.com/PaddlePaddle/PaddleOCR
- LayoutLM: github.com/microsoft/unilm
- Article on document preparation in this series
- Article on RAG fundamentals in this series


