Skip to content
BotServBotServ
n8nRAGEmbeddingsVector DatabaseSemantic SearchQdrant

RAG Workflows in n8n

Build RAG workflows in n8n. Index documents, embeddings, vector databases, semantic search and practical examples.

S

schutzgeist

7 min read
RAG Workflows in n8n

RAG Workflows in n8n

What This Article Covers

  • Building complete RAG pipelines in n8n.
  • Indexing documents, splitting them into chunks, and generating embeddings.
  • Connecting to Qdrant, Weaviate, or Pinecone.
  • How retrieval, reranking, and answer generation work.
  • Real-world examples for knowledge bases and document Q&A.

Introduction: Understanding RAG Workflows in n8n

RAG (Retrieval-Augmented Generation) flips the script: instead of asking a model to answer from memory, you first search for relevant documents and feed them as context. In n8n, you build this pipeline visually. A document comes in, gets split into chunks, embedded, and stored in a vector database. When someone asks a question, you retrieve similar chunks and pass them to the model as context.

This article is for anyone implementing RAG in n8n. For foundational concepts, see Local RAG and n8n-Ollama Integration.

Why Use RAG Workflows in n8n?

Imagine you have 500 internal documents. An employee asks, β€œHow do I file a vacation request?” Without RAG, the model makes things up. With RAG, your workflow finds the vacation policy document, feeds it as context, and the model answers correctly with sources.

RAG Workflows in n8n at a Glance

RAG in n8n consists of two pipelines: the indexing pipeline (document β†’ chunks β†’ embeddings β†’ vector database) and the query pipeline (question β†’ embedding β†’ similar chunks β†’ context β†’ answer). Both run as n8n workflows.

The core idea: search first, answer second.

Who Should Read This?

  • n8n users building RAG pipelines.
  • Knowledge managers making documents searchable.
  • Developers building Q&A systems.
  • Self-hosters running RAG locally with Ollama.

Basic familiarity with n8n and RAG concepts is expected.

Key Terms

  • RAG - Retrieval-Augmented Generation. Essential for grounding model responses in actual documents.
  • Embeddings - Vector representations of text. Used for semantic search.
  • Chunking - Breaking text into pieces. Improves retrieval quality.
  • Vector Databases - Storage for embeddings. Options include Qdrant, Weaviate, and Pinecone.
  • Ollama - Local model server. Handles embeddings and text generation.
  • Reranking - Re-ordering results by relevance. See Reranking.
  • Hybrid Search - Combines semantic and keyword matching for better hits.

Architecture

INDEXING-PIPELINE:
Dokument (PDF/DOCX)
    β”‚
    β–Ό
Text-Extraktion ──► Chunking ──► Embeddings (Ollama)
    β”‚                                β”‚
    β”‚                                β–Ό
    └────────────────────────► Vektordatenbank (Qdrant)
                                   β”‚
QUERY-PIPELINE:                      β”‚
Frage                                β”‚
    β”‚                                β”‚
    β–Ό                                β”‚
Embedding (Ollama)                   β”‚
    β”‚                                β”‚
    β–Ό                                β”‚
Vektorsuche β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
    β”‚
    β–Ό
Top-K Chunks ──► (Reranking) ──► Kontext + Frage ──► LLM (Ollama)
                                                        β”‚
                                                        β–Ό
                                                    Antwort

Setup: Qdrant + Ollama + n8n

version: "3.8"

services:
  qdrant:
    image: qdrant/qdrant:latest
    container_name: qdrant
    restart: unless-stopped
    ports:
      - "6333:6333"
    volumes:
      - qdrant_data:/qdrant/storage
    networks:
      - ai-network

  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "11434:11434"
    volumes:
      - ollama_data:/root/.ollama
    networks:
      - ai-network

  n8n:
    image: n8nio/n8n:latest
    container_name: n8n
    restart: unless-stopped
    ports:
      - "5678:5678"
    volumes:
      - n8n_data:/home/node/.n8n
    networks:
      - ai-network

volumes:
  qdrant_data:
  ollama_data:
  n8n_data:

networks:
  ai-network:
    driver: bridge

Indexing Pipeline in n8n

Step 1: Read the Document

// Read Binary File Node atau HTTP Request untuk upload
// Kemudian: Ekstrak teks

// Function Node: Ekstrak teks dari PDF
const pdfParse = require('pdf-parse');
const buffer = $input.item.binary.data;
const data = await pdfParse(Buffer.from(buffer, 'base64'));
return { text: data.text, filename: $input.item.json.filename };

Step 2: Chunk the Text

// Function Node: Bagi teks menjadi chunks
function chunkText(text, chunkSize = 500, overlap = 50) {
  const chunks = [];
  const sentences = text.split(/[.!?]\s+/);
  let current = "";

  for (const sentence of sentences) {
    if ((current + sentence).length > chunkSize) {
      chunks.push(current.trim());
      current = sentence;
    } else {
      current += " " + sentence;
    }
  }
  if (current) chunks.push(current.trim());
  return chunks;
}

const chunks = chunkText($input.item.json.text);
return chunks.map((chunk, i) => ({
  json: { text: chunk, index: i, source: $input.item.json.filename }
}));

See Chunking for different strategies.

Step 3: Generate Embeddings

// HTTP Request Node: Ollama Embeddings
{
  "method": "POST",
  "url": "http://ollama:11434/api/embeddings",
  "body": {
    "model": "nomic-embed-text",
    "prompt": "{{$json.text}}"
  }
}

See Embedding Models for model selection.

Step 4: Store in Qdrant

// HTTP Request Node: Qdrant Insert
{
  "method": "PUT",
  "url": "http://qdrant:6333/collections/dokumente/points",
  "body": {
    "points": [{
      "id": "{{$json.index}}",
      "vector": "{{$json.embedding}}",
      "payload": {
        "text": "{{$json.text}}",
        "source": "{{$json.source}}"
      }
    }]
  }
}

Query Pipeline in n8n

Step 1: Embed the Question

// HTTP Request: Query Embedding
{
  "method": "POST",
  "url": "http://ollama:11434/api/embeddings",
  "body": {
    "model": "nomic-embed-text",
    "prompt": "{{$json.question}}"
  }
}

Step 2: Find Similar Chunks

// HTTP Request: Qdrant Search
{
  "method": "POST",
  "url": "http://qdrant:6333/collections/dokumente/points/search",
  "body": {
    "vector": "{{$json.embedding}}",
    "limit": 5,
    "with_payload": true
  }
}

Step 3: Build Context and Generate Answer

// Function Node: Assemble context
const results = $input.item.json.result;
const context = results.map(r => r.payload.text).join("\n\n");
const sources = results.map(r => r.payload.source);

return {
  context,
  sources,
  question: $('Webhook').item.json.question
};

// HTTP Request: Generate answer
{
  "method": "POST",
  "url": "http://ollama:11434/api/chat",
  "body": {
    "model": "llama3.1",
    "messages": [
      {"role": "system", "content": "Beantworte die Frage basierend auf dem Kontext. Gib Quellen an. Wenn der Kontext die Frage nicht beantwortet, sage das ehrlich."},
      {"role": "user", "content": "Kontext:\n{{$json.context}}\n\nFrage: {{$json.question}}"}
    ],
    "stream": false
  }
}

Practical Example 1: Knowledge Base for FAQs

Workflow: FAQ Bot with RAG

Indexing (once / on update):
  1. Watch folder for new documents
  2. Extract text
  3. Chunking
  4. Embeddings (Ollama nomic-embed-text)
  5. Store in Qdrant

Query (on each question):
  1. Webhook receives question
  2. Generate query embedding
  3. Qdrant: retrieve top 5 chunks
  4. Feed context + question to Ollama llama3.1
  5. Return answer + sources

Practical Example 2: Indexing Email Attachments

Workflow: Make email attachments searchable

1. IMAP trigger: new email with attachment
2. Extract binary file
3. If PDF: extract text
4. Chunking + embeddings
5. Store in Qdrant with metadata (sender, date)
// Function Node: combine semantic + keyword search
const semanticResults = $input.item.json.semantic;
const keywordResults = $input.item.json.keyword;

// Reciprocal Rank Fusion
function rrf(semantic, keyword, k = 60) {
  const scores = {};
  semantic.forEach((r, i) => {
    scores[r.id] = (scores[r.id] || 0) + 1 / (k + i + 1);
  });
  keyword.forEach((r, i) => {
    scores[r.id] = (scores[r.id] || 0) + 1 / (k + i + 1);
  });
  return Object.entries(scores).sort((a, b) => b[1] - a[1]);
}

return { ranked: rrf(semanticResults, keywordResults) };

See Hybrid Search.

Improving Quality

Reranking

// After vector search: fetch top 20, then rerank
{
  "method": "POST",
  "url": "http://ollama:11434/api/chat",
  "body": {
    "model": "llama3.1",
    "messages": [
      {"role": "user", "content": "Rate the relevance of this text to the question [QUESTION] on a scale 0-10:\n\n{{$json.chunk_text}}"}
    ]
  }
}

See Reranking.

Source Attribution

// Always include sources
return {
  answer: response.message.content,
  sources: results.map(r => ({
    file: r.payload.source,
    chunk: r.payload.text.substring(0, 100) + "..."
  }))
};

See Source Attribution.

Security Considerations

  • Access Control: Who can query which documents? Filter by permissions.
  • Prompt Injection: Documents may contain injections. See Prompt Injection.
  • Data Quality: Poor documents produce poor answers. See Preparing Documents.
  • Privacy: All data stays local with Ollama and Qdrant. See Privacy.

Common Pitfalls

  • Chunks too large: Beyond 1000 characters, precision suffers. 300-800 characters is often optimal.
  • Too few chunks: Top 5 isn’t always enough. Retrieve top 20, then rerank to top 5.
  • Wrong embedding model: nomic-embed-text works well, but multilingual-e5 is better for multilingual documents.
  • Missing sources: Without sources, answers can’t be verified.
  • Hallucinations: Poor retrieval leads to model hallucinations. See Testing Quality.
  • Context too long: Too many chunks overload the context window. See Context Length.

Further Reading

Key Takeaways:

  • RAG in n8n: indexing pipeline plus query pipeline.
  • Use Ollama for embeddings (nomic-embed-text) and generation (llama3.1).
  • Qdrant as vector store, with n8n orchestrating everything.
  • Chunking strategy and reranking determine quality.
  • Always provide sources, never trust blindly.

FAQ

What is RAG?

Retrieval-Augmented Generation: instead of answering from memory, the system first retrieves relevant documents and feeds them as context to the model. This prevents hallucinations.

Which vector database?

Qdrant for self-hosting (simple, performant). Weaviate for more complex scenarios. Pinecone as a cloud option. Chroma for local development.

Which embedding model?

nomic-embed-text for German/English (via Ollama). multilingual-e5 for multilingual documents. bge-m3 for longer documents.

How large should chunks be?

300-800 characters with 50-100 character overlap. Too small loses context, too large loses precision. Test with your own documents.

How many chunks should I retrieve?

Retrieve top 20 from vector search, then rerank to top 3-5 for context. More context doesn’t always mean better answers.

How do I prevent hallucinations?

Good retrieval (finding the right chunks), clear system prompts (β€œanswer only from context”), requiring source attribution, and allowing β€œI don’t know” when uncertain.

What is Hybrid Search?

Combines semantic search (embeddings) with keyword search (BM25). Results are merged using Reciprocal Rank Fusion. Better hits than semantic alone.

What does local RAG in n8n cost?

Only hardware costs. Ollama, Qdrant, and n8n are open source. A machine with 8-16 GB VRAM is sufficient for most RAG applications.

Which documents can I index?

PDF, Word, Markdown, HTML, plaintext. Scanned documents require OCR first. Tables need specialized parsers.

How do I keep the database current?

Use a workflow with file watcher: automatically index new documents. On updates, delete old chunks and reindex. See updating your knowledge base.

Sources and Further Reading

Back to Blog
Share:

Related Posts