RAG Workflows in n8n
What This Article Covers
- Building complete RAG pipelines in n8n.
- Indexing documents, splitting them into chunks, and generating embeddings.
- Connecting to Qdrant, Weaviate, or Pinecone.
- How retrieval, reranking, and answer generation work.
- Real-world examples for knowledge bases and document Q&A.
Introduction: Understanding RAG Workflows in n8n
RAG (Retrieval-Augmented Generation) flips the script: instead of asking a model to answer from memory, you first search for relevant documents and feed them as context. In n8n, you build this pipeline visually. A document comes in, gets split into chunks, embedded, and stored in a vector database. When someone asks a question, you retrieve similar chunks and pass them to the model as context.
This article is for anyone implementing RAG in n8n. For foundational concepts, see Local RAG and n8n-Ollama Integration.
Why Use RAG Workflows in n8n?
Imagine you have 500 internal documents. An employee asks, βHow do I file a vacation request?β Without RAG, the model makes things up. With RAG, your workflow finds the vacation policy document, feeds it as context, and the model answers correctly with sources.
RAG Workflows in n8n at a Glance
RAG in n8n consists of two pipelines: the indexing pipeline (document β chunks β embeddings β vector database) and the query pipeline (question β embedding β similar chunks β context β answer). Both run as n8n workflows.
The core idea: search first, answer second.
Who Should Read This?
- n8n users building RAG pipelines.
- Knowledge managers making documents searchable.
- Developers building Q&A systems.
- Self-hosters running RAG locally with Ollama.
Basic familiarity with n8n and RAG concepts is expected.
Key Terms
- RAG - Retrieval-Augmented Generation. Essential for grounding model responses in actual documents.
- Embeddings - Vector representations of text. Used for semantic search.
- Chunking - Breaking text into pieces. Improves retrieval quality.
- Vector Databases - Storage for embeddings. Options include Qdrant, Weaviate, and Pinecone.
- Ollama - Local model server. Handles embeddings and text generation.
- Reranking - Re-ordering results by relevance. See Reranking.
- Hybrid Search - Combines semantic and keyword matching for better hits.
Architecture
INDEXING-PIPELINE:
Dokument (PDF/DOCX)
β
βΌ
Text-Extraktion βββΊ Chunking βββΊ Embeddings (Ollama)
β β
β βΌ
ββββββββββββββββββββββββββΊ Vektordatenbank (Qdrant)
β
QUERY-PIPELINE: β
Frage β
β β
βΌ β
Embedding (Ollama) β
β β
βΌ β
Vektorsuche ββββββββββββββββββββββββββ
β
βΌ
Top-K Chunks βββΊ (Reranking) βββΊ Kontext + Frage βββΊ LLM (Ollama)
β
βΌ
Antwort
Setup: Qdrant + Ollama + n8n
version: "3.8"
services:
qdrant:
image: qdrant/qdrant:latest
container_name: qdrant
restart: unless-stopped
ports:
- "6333:6333"
volumes:
- qdrant_data:/qdrant/storage
networks:
- ai-network
ollama:
image: ollama/ollama:latest
container_name: ollama
restart: unless-stopped
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
networks:
- ai-network
n8n:
image: n8nio/n8n:latest
container_name: n8n
restart: unless-stopped
ports:
- "5678:5678"
volumes:
- n8n_data:/home/node/.n8n
networks:
- ai-network
volumes:
qdrant_data:
ollama_data:
n8n_data:
networks:
ai-network:
driver: bridge
Indexing Pipeline in n8n
Step 1: Read the Document
// Read Binary File Node atau HTTP Request untuk upload
// Kemudian: Ekstrak teks
// Function Node: Ekstrak teks dari PDF
const pdfParse = require('pdf-parse');
const buffer = $input.item.binary.data;
const data = await pdfParse(Buffer.from(buffer, 'base64'));
return { text: data.text, filename: $input.item.json.filename };
Step 2: Chunk the Text
// Function Node: Bagi teks menjadi chunks
function chunkText(text, chunkSize = 500, overlap = 50) {
const chunks = [];
const sentences = text.split(/[.!?]\s+/);
let current = "";
for (const sentence of sentences) {
if ((current + sentence).length > chunkSize) {
chunks.push(current.trim());
current = sentence;
} else {
current += " " + sentence;
}
}
if (current) chunks.push(current.trim());
return chunks;
}
const chunks = chunkText($input.item.json.text);
return chunks.map((chunk, i) => ({
json: { text: chunk, index: i, source: $input.item.json.filename }
}));
See Chunking for different strategies.
Step 3: Generate Embeddings
// HTTP Request Node: Ollama Embeddings
{
"method": "POST",
"url": "http://ollama:11434/api/embeddings",
"body": {
"model": "nomic-embed-text",
"prompt": "{{$json.text}}"
}
}
See Embedding Models for model selection.
Step 4: Store in Qdrant
// HTTP Request Node: Qdrant Insert
{
"method": "PUT",
"url": "http://qdrant:6333/collections/dokumente/points",
"body": {
"points": [{
"id": "{{$json.index}}",
"vector": "{{$json.embedding}}",
"payload": {
"text": "{{$json.text}}",
"source": "{{$json.source}}"
}
}]
}
}
Query Pipeline in n8n
Step 1: Embed the Question
// HTTP Request: Query Embedding
{
"method": "POST",
"url": "http://ollama:11434/api/embeddings",
"body": {
"model": "nomic-embed-text",
"prompt": "{{$json.question}}"
}
}
Step 2: Find Similar Chunks
// HTTP Request: Qdrant Search
{
"method": "POST",
"url": "http://qdrant:6333/collections/dokumente/points/search",
"body": {
"vector": "{{$json.embedding}}",
"limit": 5,
"with_payload": true
}
}
Step 3: Build Context and Generate Answer
// Function Node: Assemble context
const results = $input.item.json.result;
const context = results.map(r => r.payload.text).join("\n\n");
const sources = results.map(r => r.payload.source);
return {
context,
sources,
question: $('Webhook').item.json.question
};
// HTTP Request: Generate answer
{
"method": "POST",
"url": "http://ollama:11434/api/chat",
"body": {
"model": "llama3.1",
"messages": [
{"role": "system", "content": "Beantworte die Frage basierend auf dem Kontext. Gib Quellen an. Wenn der Kontext die Frage nicht beantwortet, sage das ehrlich."},
{"role": "user", "content": "Kontext:\n{{$json.context}}\n\nFrage: {{$json.question}}"}
],
"stream": false
}
}
Practical Example 1: Knowledge Base for FAQs
Workflow: FAQ Bot with RAG
Indexing (once / on update):
1. Watch folder for new documents
2. Extract text
3. Chunking
4. Embeddings (Ollama nomic-embed-text)
5. Store in Qdrant
Query (on each question):
1. Webhook receives question
2. Generate query embedding
3. Qdrant: retrieve top 5 chunks
4. Feed context + question to Ollama llama3.1
5. Return answer + sources
Practical Example 2: Indexing Email Attachments
Workflow: Make email attachments searchable
1. IMAP trigger: new email with attachment
2. Extract binary file
3. If PDF: extract text
4. Chunking + embeddings
5. Store in Qdrant with metadata (sender, date)
Practical Example 3: Hybrid Search
// Function Node: combine semantic + keyword search
const semanticResults = $input.item.json.semantic;
const keywordResults = $input.item.json.keyword;
// Reciprocal Rank Fusion
function rrf(semantic, keyword, k = 60) {
const scores = {};
semantic.forEach((r, i) => {
scores[r.id] = (scores[r.id] || 0) + 1 / (k + i + 1);
});
keyword.forEach((r, i) => {
scores[r.id] = (scores[r.id] || 0) + 1 / (k + i + 1);
});
return Object.entries(scores).sort((a, b) => b[1] - a[1]);
}
return { ranked: rrf(semanticResults, keywordResults) };
See Hybrid Search.
Improving Quality
Reranking
// After vector search: fetch top 20, then rerank
{
"method": "POST",
"url": "http://ollama:11434/api/chat",
"body": {
"model": "llama3.1",
"messages": [
{"role": "user", "content": "Rate the relevance of this text to the question [QUESTION] on a scale 0-10:\n\n{{$json.chunk_text}}"}
]
}
}
See Reranking.
Source Attribution
// Always include sources
return {
answer: response.message.content,
sources: results.map(r => ({
file: r.payload.source,
chunk: r.payload.text.substring(0, 100) + "..."
}))
};
See Source Attribution.
Security Considerations
- Access Control: Who can query which documents? Filter by permissions.
- Prompt Injection: Documents may contain injections. See Prompt Injection.
- Data Quality: Poor documents produce poor answers. See Preparing Documents.
- Privacy: All data stays local with Ollama and Qdrant. See Privacy.
Common Pitfalls
- Chunks too large: Beyond 1000 characters, precision suffers. 300-800 characters is often optimal.
- Too few chunks: Top 5 isnβt always enough. Retrieve top 20, then rerank to top 5.
- Wrong embedding model: nomic-embed-text works well, but multilingual-e5 is better for multilingual documents.
- Missing sources: Without sources, answers canβt be verified.
- Hallucinations: Poor retrieval leads to model hallucinations. See Testing Quality.
- Context too long: Too many chunks overload the context window. See Context Length.
Further Reading
- Local RAG - RAG fundamentals.
- Chunking - Text splitting strategies.
- Embedding Models - Model selection.
- Vector Databases - Qdrant, Weaviate & more.
- Reranking - Improving results.
- Hybrid Search - Semantics + keywords.
- n8n-Ollama Integration - Running Ollama in n8n.
- AI Agents in n8n - Building agents with RAG.
Key Takeaways:
- RAG in n8n: indexing pipeline plus query pipeline.
- Use Ollama for embeddings (nomic-embed-text) and generation (llama3.1).
- Qdrant as vector store, with n8n orchestrating everything.
- Chunking strategy and reranking determine quality.
- Always provide sources, never trust blindly.
FAQ
What is RAG?
Which vector database?
Which embedding model?
How large should chunks be?
How many chunks should I retrieve?
How do I prevent hallucinations?
What is Hybrid Search?
What does local RAG in n8n cost?
Which documents can I index?
How do I keep the database current?
Sources and Further Reading
- Qdrant - Vector database.
- Ollama Embeddings - Embedding models.
- n8n Vector Stores - Vector store nodes.
- RAG Guide - RAG techniques.


