Skip to content
BotServBotServ
Open WebUIRAGVector DatabaseDocumentsLocal AI

Set Up RAG in Open WebUI

Enable RAG in Open WebUI. Upload documents, connect vector database, get source-based answers.

S

schutzgeist

3 min read
Set Up RAG in Open WebUI

Setting Up RAG in Open WebUI

What this article covers

  • How RAG works in Open WebUI.
  • How to upload and index documents.
  • How to configure the vector database.
  • How to access your own documents via chat.
  • Settings, optimization, and common pitfalls.

Introduction: Setting up RAG in Open WebUI

Open WebUI provides a straightforward way to run RAG locally. Users upload documents, the interface generates embeddings and stores them in a vector database. During chat, Open WebUI retrieves matching text passages and includes them in the prompt. This allows the model to answer questions based on your own documents.

RAG in Open WebUI is especially appealing for newcomers. You don’t need programming knowledge to start experimenting. For production use, it pays to understand and adjust embedding models, chunking parameters, and retrieval settings.

Why do you need RAG in Open WebUI?

Language models only know what they were trained on. Your own documents, internal processes, and specialized data are missing. RAG bridges this gap by adding relevant text passages to the prompt. This enables the model to:

  • Answer questions about contracts,
  • Explain manuals,
  • Summarize protocols,
  • reference current documents,
  • provide source citations.

RAG in Open WebUI explained

The workflow:

  1. Upload document: PDF, TXT, Markdown, or DOCX.
  2. Preparation: Open WebUI extracts text.
  3. Chunking: Text is divided into sections.
  4. Embedding: An embedding model generates vectors.
  5. Storage: Vectors are saved in the vector database.
  6. Query: During chat, matching chunks are retrieved.
  7. Answer: The model answers the question with context.

Key terminology:

  • Knowledge: In Open WebUI, a collection of uploaded documents.
  • RAG Template: A prompt that embeds the context.
  • Chunk Size: The size of text sections.
  • Chunk Overlap: Overlap between chunks.
  • Top-K: Number of the most similar chunks returned.
  • Embedding Model: Model for generating vectors.

Who is RAG in Open WebUI for?

  • Beginners wanting to test RAG without code.
  • Teams needing document-based AI quickly.
  • Small businesses with limited budgets.
  • Support teams working with FAQs.
  • Anyone wanting to stay local.

Key terminology around RAG in Open WebUI

  • Ollama: Backend for models and embeddings.
  • Chroma: Default vector database in Open WebUI.
  • BGE: Popular embedding model family.
  • nomic-embed-text: A solid open-source embedding.
  • Query Expansion: Broaden your search.
  • Reranking: Score matches retrospectively.

Step-by-step: Setting up RAG

1. Start Open WebUI

Open WebUI must be running and connected to Ollama. See BotServ.de Open WebUI.

2. Upload a document

In the chat interface, you can upload documents via the paperclip icon or in the Knowledge section. PDF, TXT, Markdown, and other formats are supported.

3. Create a Knowledge Collection

Under Admin > Knowledge, you can create collections. Documents are assigned to a collection. You can organize collections by team, topic, or project.

4. Choose an embedding model

In the settings, select your embedding model. For German text, bge-m3 or nomic-embed-text work well. The model is loaded via Ollama.

5. Chat with documents

In the chat, select your Knowledge collection. Questions will now be answered based on your documents. Sources are typically displayed.

Important settings

  • Chunk Size: Default 1000 to 1500 characters. Adjust based on your documents.
  • Chunk Overlap: Default 50 to 200 characters. Prevents context loss at boundaries.
  • Top-K: More chunks provide more context but can overwhelm the prompt.
  • RAG Template: Controls how context is embedded in the prompt.
  • Citations: Enable source attribution.

Common pitfalls

  • Wrong embedding model: English models perform poorly on German text.
  • Chunks too large: Slow responses and lost details.
  • Chunks too small: Context loss.
  • Outdated documents: Uploaded files must be refreshed.
  • No citations: Check your settings.
  • Ollama unreachable: Embedding model cannot be loaded.

Further resources

FAQ: RAG in Open WebUI

What file formats does Open WebUI support? PDF, TXT, Markdown, DOCX, and other common text formats.

Can I upload multiple documents at once? Yes, through Knowledge Collections or drag-and-drop in the chat.

How large can documents be? Depends on available storage and embedding model. Large files are automatically split into chunks.

Where are the embeddings stored? Locally in Open WebUI, by default in Chroma.

Can I restrict RAG to specific users? Yes, through Knowledge Collections and permissions.

Sources and further reading

Summary: Setting up RAG in Open WebUI

Open WebUI makes RAG accessible for beginners. Uploading documents, selecting an embedding model, and creating a Knowledge Collection is all you need to ask questions about your own content. Quality embedding models, proper chunking, source citations, and regular document maintenance are crucial. For production scenarios, fine-tuning RAG parameters pays dividends.

Back to Blog
Share:

Related Posts