Preparing Documents: Processing Data for RAG
What this article covers
- Which formats work well for RAG and what challenges each presents
- How to extract text from PDF, DOCX, HTML, and other sources
- What cleaning steps are necessary for RAG to deliver usable answers
- How metadata and normalization improve your RAG pipeline quality
- A complete working example with code you can build right away
Introduction: Document preparation explained
You’ve read about local RAG and want to get started. Your models are installed, Ollama is running, your vector database is ready. But before you can begin, there’s one task many people underestimate: preparing your documents.
This article walks you through preparing documents so your RAG application delivers clean, relevant answers. We’ll cover every common format, handle cleaning and normalization, and work through a complete example. If you haven’t yet read the RAG fundamentals, start there first.
Why do I need document preparation?
Imagine feeding raw PDF documentation into your RAG system. The PDF contains page numbers, repeated headers, tables of contents, and footnotes. You ask the question: “How do I configure the database connection?”
The system searches for similar text passages and finds, among other results, a page with the header “Chapter 4, Page 47, Version 2.1”. The answer comes back something like: “You’ll find configuration information on page 47.” That’s not a useful answer, but an artifact from unclean headers and page numbers.
This is where document preparation begins. Without it, RAG returns garbage, no matter how good your embedding model is. The quality of your answers depends directly on the quality of your data. This isn’t hyperbole, but the most important lesson you’ll learn building a RAG pipeline.
Document preparation in brief
The best analogy: document preparation is like prepping ingredients before cooking. Before you cook, you wash vegetables, chop them, remove skins, and measure out spices. Nobody throws raw, unwashed ingredients into a pot and hopes for a good meal.
RAG works the same way. You take raw documents, strip away everything distracting, restructure the text, and add metadata. Only then do the texts go into the vector database. Building the pipeline from raw document to searchable vector happens in multiple steps, which we’ll walk through in this article.
Who this article is for
This article targets beginners building their first local RAG system. You need no prior knowledge of data preparation, but should have a basic grasp of what local AI is and how RAG works. If you already know Python, you can use the code examples directly. If not, we’ll explain every step clearly.
Key concepts in document preparation
| Term | Explanation |
|---|---|
| Format | The file format of your document, such as PDF, DOCX, or Markdown. Each format requires different extraction tools. |
| Parsing | Reading and converting a document into plain text. |
| Cleaning | Removing unwanted text like headers, footers, page numbers, and boilerplate. |
| Normalization | Standardizing formatting, encoding, and whitespace for consistent data. |
| Deduplication | Removing duplicate content that skews search results. |
| Metadata | Additional information like source, date, author, and section title that help RAG contextualize answers. |
| Markdown | A simple text format with structure marked by characters like # for headings. |
| A common document format that complicates text extraction because layout and text are often mixed. | |
| HTML | Web page format that mixes text with markup tags. |
| OCR | Optical Character Recognition, the technology for extracting text from images or scanned PDFs. |
Supported formats
Not every format works equally well for RAG. Here’s an overview of common formats and their trade-offs.
PDF is the most widely used format for documentation, manuals, and reports. The downside: PDF stores layout and text together, making extraction difficult. Multi-column layouts, embedded images, and scanned pages make parsing error-prone. For scanned PDFs, you’ll need OCR, which we cover in the article on PDF and OCR.
DOCX
DOCX files from Word or LibreOffice parse well because they’re structured as XML internally. Paragraphs, headings, and lists are preserved. The drawback: formatting can be complex and not all structures convert cleanly to text.
TXT
Plain text files are the simplest format. No parsing needed, you read the text directly. The drawback: TXT offers no structure, so you must manually identify headings and sections.
Markdown
Markdown is the best format for RAG because it’s both structured and simple. Headings, lists, and code blocks are marked directly in the text. If you create documents yourself, use Markdown.
HTML
HTML from websites contains lots of boilerplate: navigation, footers, ads, scripts. You need to strip the markup and extract only the main content. Tools like BeautifulSoup work well but require effort.
CSV
CSV suits tabular data. Each row becomes a text chunk, column headers become metadata. The drawback: tables don’t translate well to flowing text, and RAG models struggle with structured data.
Step 1: Collect and organize your documents
Before you parse a single document, gather all files in one location and give them a consistent structure. This sounds obvious but saves hours later.
Put all source documents in one folder and name them consistently. A good naming scheme is date_topic_version.pdf, like 2026-01-15_api-docs_v2.pdf. At a glance, you’ll see which version you have and when it was created.
For each document, decide what metadata you want to capture. A simple schema looks like this:
metadata_schema = {
"source": "filename.pdf",
"title": "Document Title",
"author": "Author or Organization",
"date": "2026-01-15",
"section": "Section or Chapter",
"page": 42
}
You’ll use this schema for every text chunk you extract. Consistency here pays off when your RAG application provides source citations, as described in the article on source attribution.
Step 2: Extract Text
Each format needs the right tools. Here are the most important ones with brief code examples.
PDF with PyPDF2
from pypdf import PdfReader
reader = PdfReader("dokument.pdf")
text = ""
for page in reader.pages:
text += page.extract_text() + "\n"
PyPDF2 pulls out plain text but ignores layout information. With multi-column PDFs, the order of text chunks can get jumbled. More sophisticated tools like pdfplumber or unstructured give better results.
DOCX with python-docx
from docx import Document
doc = Document("dokument.docx")
text = ""
for paragraph in doc.paragraphs:
text += paragraph.text + "\n"
python-docx reads paragraphs sequentially and preserves order. Identify headings by the style attribute, like paragraph.style.name == "Heading 1".
HTML with BeautifulSoup
from bs4 import BeautifulSoup
with open("seite.html", "r", encoding="utf-8") as f:
soup = BeautifulSoup(f, "html.parser")
# Remove scripts and styles
for tag in soup(["script", "style", "nav", "footer"]):
tag.decompose()
text = soup.get_text(separator="\n")
BeautifulSoup strips unwanted tags and extracts visible text. Tailor the list of tags to remove based on your sources.
Markdown and TXT
with open("dokument.md", "r", encoding="utf-8") as f:
text = f.read()
Read Markdown and TXT files directly, no special tool required.
Step 3: Clean the Text
After extraction, your text contains plenty of noise that confuses RAG. Cleaning is the most important step for quality data.
Remove headers and footers
PDFs repeat headers and footers on every page, like document titles or version numbers. Strip these patterns with regex:
import re
# Remove page numbers like "Page 12" or "- 12 -"
text = re.sub(r"(Seite|Page)\s*\d+", "", text)
text = re.sub(r"-\s*\d+\s*-", "", text)
# Remove repeated headers (adjust for your documents)
text = re.sub(r"Kapitel \d+", "", text)
You’ll need to adapt these patterns to your specific documents. There’s no universal solution because every document has different headers.
Fix encoding issues
Wrong encoding produces garbled characters like ü instead of ü. Convert everything to UTF-8:
text = text.encode("utf-8", errors="ignore").decode("utf-8")
Remove boilerplate
Tables of contents, copyright notices, and disclaimers repeat across every page and waste space in your vector database. Strip them out:
boilerplate_patterns = [
r"Inhaltsverzeichnis.*?(?=\n\n)",
r"© \d{4}.*",
r"Alle Rechte vorbehalten.*",
]
for pattern in boilerplate_patterns:
text = re.sub(pattern, "", text, flags=re.DOTALL)
Normalize whitespace and blank lines
# Collapse multiple spaces into one
text = re.sub(r"[ \t]+", " ", text)
# Collapse multiple blank lines into one
text = re.sub(r"\n{3,}", "\n\n", text)
# Remove leading and trailing whitespace
text = text.strip()
Step 4: Add Metadata
Metadata matters more for RAG than most people realize. Without it, your system doesn’t know where a text chunk came from or how fresh it is. This leads to answers containing outdated information from old versions.
For each text chunk, capture at least these fields:
- source: Filename or URL of the original document
- title: Document or section title
- date: Creation or modification date
- author: Author or publishing organization
- section: Chapter or section heading
- page: Page number for PDFs
In Python, store metadata as a dictionary alongside the text:
chunk = {
"text": "The extracted text chunk...",
"metadata": {
"source": "handbuch_v2.pdf",
"title": "Configure Database Connection",
"date": "2026-01-15",
"author": "BotServ",
"section": "Chapter 4",
"page": 47
}
}
Later, when you search, you can filter by metadata, like only current versions or specific chapters. This makes your RAG application sharper and results traceable.
Step 5: Normalization
Normalization ensures all your documents follow the same format regardless of source. This simplifies chunking and later retrieval.
Key normalization steps:
- Standardize encoding: Convert everything to UTF-8
- Standardize line endings: Convert
\r\nand\rto\n - Normalize whitespace: Reduce multiple spaces and blank lines
- Clean special characters: Remove invisible characters like zero-width spaces
def normalize_text(text):
# Standardize line endings
text = text.replace("\r\n", "\n").replace("\r", "\n")
# Remove zero-width spaces and similar invisible characters
text = re.sub(r"[\u200b\u200c\u200d\ufeff]", "", text)
# Normalize whitespace
text = re.sub(r"[ \t]+", " ", text)
text = re.sub(r"\n{3,}", "\n\n", text)
return text.strip()
Apply this function to every extracted text before further processing. Consistency here pays off when you later combine different document types in a vector database.
Best Practices for Clean RAG Data
Here are the key tips that make a difference in practice:
- Start small: Begin with 5 to 10 documents, not 500. You’ll catch problems early.
- Manually review extracted text: Read through the first few extracts yourself. If they’re unreadable to you, they’re useless to RAG.
- Use Markdown as an intermediate format: Convert everything to Markdown before loading into your vector database. This preserves structure and lets you track changes.
- Document your pipeline: Write down which steps you run and in what order. When something breaks, you can isolate the issue.
- Version your data: When you update documents, keep old versions. Then you can compare whether new data produces better answers.
- Remove duplicates: Identical content across multiple documents skews search results. Check for duplicates before loading.
- Test with real questions: Ask your RAG application questions actual users would ask. This shows whether your preparation is adequate.
Example: Preparing a PDF Manual
Let’s walk through a complete example. You have a 50-page PDF manual and want to prepare it for RAG.
Step 1: Load the PDF and extract text
from pypdf import PdfReader
import re
reader = PdfReader("handbuch.pdf")
pages = []
for i, page in enumerate(reader.pages):
text = page.extract_text()
pages.append({"text": text, "page": i + 1})
Step 2: Clean the text
def clean_page(text):
# Remove page numbers
text = re.sub(r"-\s*\d+\s*-", "", text)
text = re.sub(r"Seite\s*\d+", "", text)
# Remove repeated headers (customize!)
text = re.sub(r"Handbuch v\d\.\d", "", text)
# Normalize whitespace
text = re.sub(r"[ \t]+", " ", text)
text = re.sub(r"\n{3,}", "\n\n", text)
return text.strip()
for page in pages:
page["text"] = clean_page(page["text"])
Step 3: Add metadata
for page in pages:
page["metadata"] = {
"source": "handbuch.pdf",
"title": "System Manual",
"date": "2026-01-15",
"author": "BotServ",
"page": page["page"]
}
Step 4: Split into chunks
Here we use simple paragraph-based splitting. For production systems, use dedicated chunking strategies as described in the chunking article.
chunks = []
for page in pages:
paragraphs = page["text"].split("\n\n")
for para in paragraphs:
if len(para.strip()) > 50: # Skip very short paragraphs
chunks.append({
"text": para.strip(),
"metadata": page["metadata"]
})
print(f"{len(chunks)} chunks created")
Step 5: Generate embeddings
Now convert your cleaned chunks to vectors. Which models work best is covered in the embedding models article.
# Pseudocode, see embedding models article for details
for chunk in chunks:
chunk["embedding"] = embed(chunk["text"])
This example shows the complete workflow. In practice, you’ll adapt the cleaning patterns to your specific documents, but the structure stays the same.
Common Pitfalls in Document Preparation
- Ignoring headers and footers: Many developers overlook the fact that repeated headers can skew search results. Remove them consistently.
- Failing to check encoding: If you don’t test for special characters, corrupted text ends up in your database and degrades embedding quality.
- Chunks that are too large: When a chunk exceeds reasonable length, the embedding model loses focus. Keep chunks between 200 and 800 words.
- Omitting metadata: Without metadata, you can’t cite sources or apply filters. This makes RAG impractical for production use.
- Not removing duplicates: If the same text appears multiple times in your database, it will show up in every response. That’s not useful.
- Including tables of contents: These contain only page references and no actual content. They waste space and distort search results.
- Treating tables as plain text: Tables lose their structure during parsing. Consider whether to represent them as structured data instead of text.
- Never testing with real questions: If you only build the pipeline but never test with actual user queries, you won’t discover gaps in your data.
Hardware, Cost, and Security in Document Preparation
Document preparation itself is computationally cheap. Text extraction and cleaning run fine on a standard laptop. You don’t need a GPU for these steps.
Costs mainly come from storage for processed data and from the embedding calculation that happens next. If you use local embedding models, you avoid API charges entirely.
Security matters particularly when processing sensitive documents. Since everything runs locally on your machine, your data stays with you. That’s one of the biggest advantages of local AI. Still, make sure intermediate files don’t sit around unencrypted, especially if you’re storing metadata with personal information.
Further Reading and Resources on Document Preparation
- Local RAG - Overview of all RAG articles
- RAG Fundamentals - How RAG works
- Chunking - Splitting text correctly
- PDF and OCR - Processing scanned PDFs
- Embedding Models - Generating vectors from text
- Source Attribution - Making answers traceable
- Ollama - Running local models
FAQ: Preparing Documents, Common Questions
Do I have to clean every document manually?
No, automate cleaning with scripts. Once written, they run over all documents. You only spot-check to verify quality.
What format works best for RAG?
Markdown is ideal because it combines structure with simplicity. If you’re creating documents from scratch, use Markdown. For existing documents, convert them to Markdown as an intermediate format.
Do I need OCR for all PDFs?
Only for scanned PDFs that contain text as images. Digital PDFs exported from word processors already contain selectable text and don’t need OCR.
How large should my chunks be?
A good guideline is 200 to 800 words per chunk. Chunks that are too small lose context, and those that are too large lose focus. Test different sizes with real questions and compare results.
How should I handle tables in documents?
Tables don’t translate well to plain text. You can preserve them as Markdown tables or store them as structured metadata. Some RAG systems handle structured data well, others don’t. Test what works best for your use case.
How do I handle multiple languages?
If your documents are multilingual, use an embedding model that supports multiple languages. Clean text language-specifically, using different patterns for headers and footers. You can also run language detection first and separate documents by language.
Can I update documents later?
Yes, and you should. When a document changes, remove the old chunks from your vector database and load the new ones. Metadata helps you locate and replace old versions precisely.
How many documents do I need for a useful RAG system?
It depends on your use case. For specific documentation, 10 to 50 documents suffice. For a broad knowledge base, you’ll need more. Quality of prepared data matters more than quantity.
What’s the difference between cleaning and normalization?
Cleaning removes unwanted content like headers and boilerplate. Normalization standardizes format, such as encoding and whitespace. Both steps are necessary and complement each other.
Do I need programming skills for document preparation?
For automated pipelines, yes, you need basic Python knowledge. For small amounts, you can clean documents manually, but that doesn’t scale. The code examples in this article are written so you can follow them even as a beginner.
Sources and Further Reading
- PyPDF2 documentation: github.com/py-pdf/pypdf
- python-docx documentation: python-docx.readthedocs.io
- BeautifulSoup documentation: crummy.com/software/BeautifulSoup
- LangChain documentation for document processing: python.langchain.com
- Unstructured Library: github.com/Unstructured-IO/unstructured
- Article on RAG Fundamentals in this series
- Article on Chunking in this series


