Keeping Your Knowledge Base Current: RAG Data Maintenance
What this article covers
- Why a RAG knowledge base needs regular updates and what happens if you skip them.
- The two main strategies: full rebuild and incremental updates.
- Step-by-step guidance for new, modified, and outdated documents.
- Versioning and automation to keep your knowledge base current.
- Common pitfalls and how to avoid them.
Introduction: Knowledge base maintenance explained
A RAG system is only as good as its data. You’ve collected documents, split them into chunks, processed them with an embedding model, and stored them in a vector database. So far, so good. But documents change. New ones arrive, old ones become invalid, some get revised.
This is where knowledge base maintenance begins. This article shows you how to keep your RAG knowledge base current without starting from scratch each time. You’ll learn what strategies exist, when each one makes sense, and how to automate the process.
If you haven’t yet covered the RAG fundamentals, read that first. It explains how the individual components work together.
Why do you need updates?
Imagine your company has an employee handbook. It covers vacation policies, work hours, and guidelines. You loaded this handbook into your RAG system six months ago. Now someone asks, “How many vacation days do I get per year?”
The RAG system searches the vector database, finds the old entry, and answers, “30 days.” The problem: the handbook was updated two months ago. It’s now 28 days. Your RAG system doesn’t know this because nobody removed the old version and added the new one.
This isn’t theoretical. It happens anywhere documents aren’t static. Contracts change, regulations get updated, product manuals get new versions, internal policies get revised. A RAG system without maintenance becomes unreliable over time.
Other examples of stale data:
- A product catalog where items have been removed but the system still recommends them.
- An outdated privacy policy that no longer matches current legal requirements.
- Expired offers that the system still presents as valid.
Knowledge base maintenance in a nutshell
Think of a library. New books arrive regularly. Some books get replaced by newer editions. Others are discarded because they’re outdated. The librarian keeps track of what’s new, what’s been replaced, and what’s been removed. Without this maintenance, the library becomes unusable because visitors find outdated information.
Your RAG knowledge base works the same way. The vector database is the library, the documents are the books, and you’re the librarian. New documents must be added, changed ones replaced, and outdated ones removed. This process is called knowledge base maintenance.
Who this article is for
This article is for beginners who already have a working RAG system and want to learn how to maintain it long term. You don’t need deep programming knowledge, but you should be familiar with Local RAG and Document preparation.
If you’re asking yourself what local AI is in the first place, the article What is local AI? can help.
Key terminology for maintenance
| Term | Meaning |
|---|---|
| Knowledge base | The collection of all documents and chunks in your vector database |
| Collection | A named space in the vector database containing related data |
| Upsert | Update an entry if it exists, otherwise create it new (Update + Insert) |
| Delete | Remove an entry or entries from the vector database |
| Versioning | Tracking which version of a document is stored |
| Incremental Update | Update only changed or new documents, not everything |
| Full Rebuild | Complete reconstruction of the vector database, all data re-embedded |
| Stale | An outdated entry that no longer matches the current state of the source document |
| Chunk ID | Unique identifier for a single chunk in the database |
| Embedding | A numerical vector representing the meaning of a text passage |
What happens with stale data?
Stale data isn’t just inelegant, it actively damages your RAG system. Here are the most common consequences:
Incorrect answers: The system answers based on old information. This might harmlessly mean an outdated product date is mentioned. In worse cases, the system gives false security or legal information.
Contradictory answers: If you have both an old and new version of the same document in the database, the system might deliver the old answer to one query and the new answer to another. This looks unreliable and confusing.
Compliance issues: In regulated fields like finance, medicine, or data protection, it’s critical that only current documents are used. Outdated policies can have legal consequences.
Eroding trust: When users notice answers are outdated, they lose confidence in the system. That’s hard to rebuild.
Strategy 1: Full Rebuild
A full rebuild is the simplest approach. You delete all data from the vector database and reload all documents. Every document gets split into chunks again, every chunk gets re-embedded and stored.
When does this make sense?
- When first building the system.
- If your chunking strategy or embedding model has changed.
- If the database is corrupted or inconsistent.
- For small knowledge bases that rebuild quickly anyway.
- If you have no tracking system and don’t know what changed.
Advantages:
- Simple to implement, no change tracking needed.
- Guarantees a consistent state afterward.
- No stale entries remain.
Disadvantages:
- Takes a long time for large knowledge bases.
- Every document needs re-embedding, which costs compute.
- The system is unavailable during the rebuild.
A full rebuild is like hitting your database with a fresh start. For small projects or infrequent maintenance windows, it’s completely sufficient.
Strategy 2: Incremental Updates
Incremental updates are the more efficient approach for larger knowledge bases. Instead of reloading everything, you update only what changed. New documents get added, modified ones get replaced, removed ones get deleted.
When does this make sense?
- For large knowledge bases where a full rebuild would take too long.
- When documents change frequently but only in small parts.
- When you have a change detection system in place.
- When the system needs to stay available during updates.
Advantages:
- Much faster than a full rebuild.
- Lower compute cost, only changed documents get re-embedded.
- System stays usable.
Disadvantages:
- More complex to implement.
- Requires tracking which version of each document is stored.
- Detection errors lead to stale entries.
Incremental updates are the strategy of choice for production systems that need continuous maintenance.
Step 1: Adding New Documents
Adding new documents follows the same process as your initial setup. You’ll go through the familiar steps from Preparing Documents:
- Prepare: Load the document, clean it, and verify its content. Remove unnecessary formatting and ensure the text is readable.
- Chunk: Split the document into sections as described in Chunking.
- Embed: Convert each chunk into a vector using your embedding model.
- Store: Insert the vectors along with metadata into your vector database.
Including metadata is essential. At minimum, save these fields:
source: Path or URL of the original document.version: Version number or date of the document.chunk_id: Unique identifier for the chunk.created_at: Timestamp when the chunk was added.
With this metadata, you can later update or delete documents precisely.
Here’s a simple example using Chroma:
import chromadb
client = chromadb.PersistentClient(path="./vectordb")
collection = client.get_or_create_collection("wissen")
collection.add(
documents=["Urlaubsregelung 2026: 28 Tage pro Jahr"],
metadatas=[{"source": "handbuch.pdf", "version": "2026-08", "chunk_id": "handbuch_001"}],
ids=["handbuch_001"]
)
Step 2: Updating Changed Documents
When a document changes, simply adding the new version isn’t enough. You’d end up with both old and new chunks in the database simultaneously, leading to contradictory answers.
Follow this workflow:
- Detect the change: Compare the current document version against what’s stored. Hashes or version numbers work well for this.
- Delete old chunks: Remove all chunks belonging to the old version. Find them using metadata like
sourceandversion. - Add new chunks: Process the updated document, embed it, and store it with the new version number.
Many vector databases support upsert, which simultaneously updates and inserts records. With Chroma, it looks like this:
# Delete old chunks for this document
collection.delete(
where={"source": "handbuch.pdf"}
)
# Add the new version
collection.add(
documents=["Neue Urlaubsregelung 2026: 28 Tage pro Jahr"],
metadatas=[{"source": "handbuch.pdf", "version": "2026-09", "chunk_id": "handbuch_001"}],
ids=["handbuch_001"]
)
Alternatively, use upsert if your database supports it. Upsert overwrites existing IDs and adds new ones.
Step 3: Removing Obsolete Documents
Sometimes documents become completely invalid. A product gets discontinued, a policy is withdrawn, a contract expires. In these cases, you need to remove the corresponding chunks from the database.
You can delete in several ways:
By source: Remove all chunks belonging to a specific document.
collection.delete(where={"source": "altes_handbuch.pdf"})
By date: Remove all chunks added before a certain date.
collection.delete(where={"created_at": {"$lt": "2026-01-01"}})
By metadata: Remove all chunks with a specific property, such as an outdated version.
collection.delete(where={"version": "2025-12"})
Important: Don’t blindly delete everything that’s old. Many documents remain valid for years. Verify that a document is truly obsolete before deleting it.
Versioning and Traceability
Versioning is the key to clean updates. Without it, you don’t know which version of a document is in the database, and you can’t update it precisely.
Simple versioning stores the following metadata with each chunk:
source: Where does the document come from?version: Which version is it? This can be a number, date, or hash value.updated_at: When was it last updated?
You can query which versions exist in the database at any time:
results = collection.get(
where={"source": "handbuch.pdf"},
include=["metadatas"]
)
for item in results["metadatas"]:
print(item["version"], item["updated_at"])
Advanced versioning uses hash values. Calculate a hash of the document’s content and store it. When content changes, the hash changes too. This lets you detect changes reliably without manual comparison.
Automation
Manual updates work fine for small knowledge bases. But once you have dozens or hundreds of documents, it becomes tedious. Automation helps here.
Cron jobs: Set up a regular job that monitors your document folder. The job compares current files against stored versions and updates the database as needed.
Watch folder: A script monitors a folder for changes. Whenever a file is added, modified, or deleted, the database updates automatically.
Here’s a simple example of a cron job in Python:
import os
import hashlib
import chromadb
def datei_hash(pfad):
with open(pfad, "rb") as f:
return hashlib.md5(f.read()).hexdigest()
def aktualisiere_wissensbestand(ordner, collection):
for dateiname in os.listdir(ordner):
pfad = os.path.join(ordner, dateiname)
if not os.path.isfile(pfad):
continue
hash_wert = datei_hash(pfad)
# Pruefen, ob sich etwas geaendert hat
bestehend = collection.get(
where={"source": dateiname},
include=["metadatas"]
)
if bestehend["metadatas"]:
aktueller_hash = bestehend["metadatas"][0].get("hash")
if aktueller_hash == hash_wert:
continue # Keine Aenderung
# Alte Version loeschen
collection.delete(where={"source": dateiname})
# Neue Version hinzufuegen
with open(pfad, "r", encoding="utf-8") as f:
text = f.read()
collection.add(
documents=[text],
metadatas=[{"source": dateiname, "hash": hash_wert, "version": hash_wert[:8]}],
ids=[dateiname]
)
client = chromadb.PersistentClient(path="./vectordb")
collection = client.get_or_create_collection("wissen")
aktualisiere_wissensbestand("./dokumente", collection)
This script is a starting point. For production use, extend it with error handling, logging, and proper chunking.
Common Pitfalls When Updating
1. No metadata stored: Without metadata, you can’t update or delete selectively. You’ll need a full rebuild instead. Always store at least source and version.
2. Old versions not deleted: Adding a new version without removing the old one creates duplicates with conflicting content. This produces contradictory answers.
3. Changes not detected: Without a change detection system, you either update too frequently (wasted effort) or too seldom (stale data). Hashes or version numbers solve this.
4. Embedding model changed, data not rebuilt: If you switch to a different embedding model, old vectors are no longer compatible. A full rebuild is necessary, otherwise search results become unreliable.
5. Chunking changed, IDs not adjusted: When you change your chunking strategy, chunk boundaries shift. Old chunk IDs no longer match new chunks. Delete the old chunks and add the new ones.
6. No backup before updating: If something goes wrong during an update, your data could be lost. Back up your vector database beforehand.
7. Concurrent writes: Multiple processes writing to the database simultaneously can cause conflicts. Lock the database during updates or use a queue.
8. Forgotten documents in subdirectories: If your script only searches the root folder, documents in subdirectories get missed. Make sure to search recursively.
Hardware, Costs, and Security During Updates
Hardware: Updates are primarily compute-intensive for embedding. Processing many documents at once can saturate your CPU or GPU. Incremental updates have minimal overhead since only a few documents get re-embedded. A full rebuild on large collections can take several hours.
Costs: Running locally incurs no direct costs beyond electricity and hardware wear. If you use cloud-based embeddings, each API call costs money. Incremental updates are significantly cheaper than a full rebuild in this scenario.
Security: Your data stays on your machine as long as you run the embedding model and vector database locally. This is one of the main advantages of local AI. Make sure backups are encrypted, especially if they contain sensitive documents.
Further Resources and Information on Updates
- Local RAG - Overview of all RAG articles.
- RAG Fundamentals - How the building blocks work together.
- Preparing Documents - Documents must be clean before you can update them.
- Chunking - How you split documents affects how you update them.
- Embedding Models - Switching models requires a rebuild.
- Vector Databases - Your database is where everything is stored.
- Testing RAG Quality - Check quality after each update.
FAQ: Updating Your Knowledge Base - Common Questions
How often should I update my knowledge base?
It depends on how frequently your documents change. With daily content updates, daily refreshes make sense. Stable documents might only need weekly or monthly updates.
Full rebuild or incremental updates, which is better?
Full rebuilds are simpler for small knowledge bases or when switching embedding models or chunking strategies. Incremental updates are more efficient for large, actively maintained collections.
What happens if I don’t delete old chunks?
Your database will contain both old and new versions at the same time. The RAG system may return contradictory answers because it doesn’t know which version is current.
How do I know if a document has changed?
The simplest approach is a hash of the file contents. If the content changes, so does the hash. Alternatively, you can track version numbers or file modification timestamps.
Do I need to restart the vector database after each update?
No, most vector databases support updates while running. For large full rebuilds, you might want to briefly lock the database to prevent conflicts.
Can I automate updates?
Yes, using Cron jobs or scripts that monitor a directory. The database updates automatically when files change. This is recommended for production systems.
What do I do if I switch embedding models?
You must perform a full rebuild. The old vectors are incompatible with the new model because each model produces vectors with different meanings and dimensions.
How do I handle deleted documents?
Delete all chunks belonging to the deleted document’s source. Use metadata fields, typically source, to identify and remove the relevant entries.
Do I need a backup before updating?
Absolutely. If something goes wrong, you can restore the previous state. Backups are essential before full rebuilds.
How do I verify that an update was successful?
Ask test questions you know the answers to and check whether the system returns current information. More details in Testing RAG Quality.
What are metadata and why do they matter?
Metadata is additional information attached to each chunk, such as source, version, or date. It lets you update and delete selectively without rebuilding the entire database.
Sources and Further Reading
- Chroma Documentation: Collections and Metadata
- Qdrant Documentation: Points and Filtering
- LangChain Documentation: Document Loaders and Vector Stores
- LlamaIndex Documentation: Index Management


