Chunking
Introduction
Long documents don’t fit into a language model in their entirety. Before processing, you need to break them down into smaller sections. This step is called chunking. Well-executed chunking is one of the most important levers for improving the quality of a RAG pipeline.
Chunking - In A Nutshell
Chunking means splitting a document into sections of a specific length. Each section is later converted into an embedding vector and stored in a vector database. The search then retrieves not the entire document, but the matching section.
What matters is that each section makes sense on its own. Text cut off mid-sentence produces poor results. A good chunk contains a complete thought, a unit of information, or a paragraph that stands alone.
Key Terms and Components
| Term | Meaning |
|---|---|
| Chunk | A single text section |
| Chunk Size | Maximum size of a chunk in characters or tokens |
| Chunk Overlap | Number of characters or tokens shared by two adjacent chunks |
| Token | Processing unit of the language model |
| Semantic Chunking | Division by meaning rather than fixed size |
Practical Relevance
Why Is Chunking So Important?
Poor chunking leads to poor answers. If a section ends in the middle of important context, the model loses critical information. If a chunk is too long, the search gets lost in generic passages. If it’s too short, you lose connections between ideas.
A contract is a good example. If a chunk ends at the conclusion of section 2.1 and section 2.2 begins in the next chunk, the model might miss the relationship between them. An overlap or division by section helps here.
Common Chunking Strategies
| Strategy | Description | When to Use |
|---|---|---|
| Fixed Size | Every chunk has a set length | Quick first version |
| Overlap | Chunks share a few sentences | Preserve context transitions |
| By Paragraph | Division according to paragraphs | Structured documents |
| Semantic | Division by meaning boundaries | High quality, more effort |
| Recursive | Start with large units, split as needed | Flexible adaptation |
Example: Python with LangChain
from langchain.text_splitter import RecursiveCharacterTextSplitter
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=50,
separators=["\n\n", "\n", ". ", " ", ""]
)
chunks = text_splitter.split_text(document)
This example splits text into sections of roughly 500 characters. It prefers to split at paragraph boundaries, line breaks, and sentence endings. The 50-character overlap ensures that context isn’t lost at chunk boundaries.
Choosing the Right Size
A typical chunk size falls between 200 and 1,000 tokens. Smaller chunks are more precise during search but provide less context. Larger chunks offer more context but can include irrelevant information.
For German text with longer compound words, 400 to 800 characters often works well. For technical documentation or legal texts, structured chunks based on headings tend to perform better.
Hardware, Cost, and Security Considerations
Chunking itself is resource-efficient. It affects vector database storage because you store more small chunks. The number of entries grows while the embedding size per entry stays the same. Costs are minimal. From a security perspective, everything stays local as long as embeddings and the database remain on your own machine.
More AI Information and Topics
- Chunking divides documents into meaningful sections.
- Overlaps preserve connections at chunk boundaries.
- Structured chunking by paragraph often outperforms pure length-based division.
- The right chunk size depends on document type.
Learn more about embeddings in Embedding Models and vector databases in Vector Databases.
FAQ - Common Questions About Chunking
What is the best chunk size?
There’s no universally optimal size. For German text, 400 to 800 tokens often works well. Testing with your own questions is essential.
Should I always use overlap?
Yes, a small overlap of 10 to 20 percent helps prevent losing context at section boundaries.
Can I create chunks based on headings?
Yes, that often works better than pure length-based division. Markdown or HTML structures are particularly suited for this.
What happens with chunks that are too large?
Search may return long, generic sections. The model receives too much context and answers become less precise.
Tools and Further Reading
LangChain, LlamaIndex, and unstructured offer robust chunking capabilities. For simple cases, Python scripts with regular expressions are sufficient.
Sources
- LangChain Text Splitter
- LlamaIndex Chunking Documentation
- Unstructured Library


