Skip to content
BotServBotServ
n8nOllamaIntegrationAI WorkflowLocal AI

Integrate n8n with Ollama

Integrate n8n with Ollama using HTTP requests, AI workflows, classification, summarization and practical examples.

S

schutzgeist

8 min read
Integrate n8n with Ollama

Integrating n8n with Ollama

What this article covers

  • How to connect n8n with Ollama.
  • Configuring HTTP Request nodes for the Ollama API.
  • Practical examples: classification, summarization, translation, and extraction.
  • Building RAG workflows with Ollama embeddings.
  • Best practices for error handling, caching, and performance.

Introduction: understanding n8n and Ollama integration

n8n is a workflow automation platform that connects processes visually. Ollama is a local AI model server. When you combine them, you can build workflows that leverage AI locally: classify emails, summarize documents, translate text, all without cloud services, without API costs, and with complete control over your data.

This article is for users who work with n8n for workflow automation and want to integrate Ollama for local AI. You should understand how n8n works and how Ollama operates.

Why do you need n8n with Ollama?

Imagine you want to automatically classify emails. Without AI, you need rules: “If subject contains X, move to folder Y”. But rules are rigid. With Ollama in n8n, the AI understands context and classifies accordingly. Everything stays local, and email content never leaves your network.

How n8n and Ollama work together

n8n sends HTTP requests to the Ollama API. The response flows back into the workflow for further processing. Everything runs locally: n8n on your server, Ollama on your server, no data leaves your network.

The core principle is straightforward: n8n controls the workflow, Ollama provides the AI.

Who this article is for

  • n8n users integrating AI into workflows.
  • Ollama users automating workflows.
  • Self-hosters running local AI.
  • Automation engineers solving repetitive tasks with AI.

Prior experience with n8n and Ollama is assumed.

Key terms

  • n8n - Workflow automation. Use when: you need a visual workflow tool.
  • Ollama - Local model server. Use when: you need an AI backend.
  • HTTP Request Node - n8n node for API calls. Use when: calling Ollama.
  • Ollama REST API - Ollama interface. Use when: n8n communicates with Ollama.
  • Embeddings - Vector representations. Use when: building RAG workflows.
  • RAG - Retrieval-Augmented Generation. Use when: building knowledge workflows.
  • Docker - Containers. Use when: running n8n and Ollama.

Setup: connecting n8n and Ollama

1. Ollama is running

# Ollama runs on localhost:11434
ollama serve
ollama list  # Check if models are loaded

2. n8n is running

# Start n8n with Docker
docker run -d \
  --name n8n \
  -p 5678:5678 \
  -v n8n_data:/home/node/.n8n \
  --restart unless-stopped \
  n8nio/n8n:latest

3. Networking

If both run in Docker, they must be on the same network:

docker network create n8n-ollama

docker run -d --name ollama --network n8n-ollama -p 11434:11434 ollama/ollama
docker run -d --name n8n --network n8n-ollama -p 5678:5678 n8nio/n8n:latest

# In n8n: URL is http://ollama:11434 instead of http://localhost:11434

See Docker networking for details.

Calling the Ollama API in n8n

Basic HTTP Request

// n8n HTTP Request Node configuration:
{
  "method": "POST",
  "url": "http://ollama:11434/api/chat",  // or http://localhost:11434
  "headers": {
    "Content-Type": "application/json"
  },
  "body": {
    "model": "llama3.1",
    "messages": [
      {"role": "user", "content": "{{$json.prompt}}"}
    ],
    "stream": false
  }
}

Function Node for Ollama

// In n8n Function Node:
const response = await this.helpers.httpRequest({
  method: 'POST',
  url: 'http://ollama:11434/api/chat',
  body: {
    model: 'llama3.1',
    messages: [
      { role: 'system', content: 'You are a helpful assistant.' },
      { role: 'user', content: $input.item.json.prompt }
    ],
    stream: false
  },
  json: true
});

return { response: response.message.content };

Practical example 1: email classification

// Workflow:
// 1. IMAP trigger: New email
// 2. Function: Build prompt
// 3. HTTP Request: Ollama classifies
// 4. Switch: Based on category
// 5. Various actions

// Function Node (build prompt):
const email = $input.item.json;
return {
  prompt: `Classify this email into one of these categories: support, sales, billing, spam, other.

Subject: ${email.subject}
From: ${email.from}
Body: ${email.body}

Reply with only the category, nothing else.`
};

// HTTP Request Node:
{
  "method": "POST",
  "url": "http://ollama:11434/api/chat",
  "body": {
    "model": "llama3.1",
    "messages": [{"role": "user", "content": "{{$json.prompt}}"}],
    "stream": false
  }
}

// Function Node (process result):
const category = $input.item.json.message.content.trim().toLowerCase();
return { category, email: $input.item.json.email };

See email automation for details.

Practical example 2: document summarization

// Workflow:
// 1. Watch: New folder
// 2. Read Binary File: Read PDF
// 3. Function: Extract text
// 4. HTTP Request: Ollama summarizes
// 5. Write File: Save summary

// Function Node:
const document = $input.item.json;
return {
  prompt: `Summarize the following document in 5 sentences:\n\n${document.text}`
};

Practical example 3: translation

// Function Node:
return {
  prompt: `Translate to English:\n\n${$input.item.json.text}`
};

// HTTP Request Node:
{
  "method": "POST",
  "url": "http://ollama:11434/api/chat",
  "body": {
    "model": "qwen2.5",  // Good model for translation
    "messages": [{"role": "user", "content": "{{$json.prompt}}"}],
    "stream": false
  }
}

Practical example 4: RAG with Ollama embeddings

// 1. Create embeddings for documents
const embedding = await this.helpers.httpRequest({
  method: 'POST',
  url: 'http://ollama:11434/api/embeddings',
  body: {
    model: 'nomic-embed-text',
    prompt: document.text
  },
  json: true
});

// 2. Store in vector database (e.g. Qdrant, Pinecone)
await this.helpers.httpRequest({
  method: 'PUT',
  url: 'http://qdrant:6333/collections/docs/points',
  body: {
    points: [{
      id: docId,
      vector: embedding.embedding,
      payload: { text: document.text }
    }]
  },
  json: true
});

// 3. On question: Find similar documents
const queryEmbedding = await this.helpers.httpRequest({
  method: 'POST',
  url: 'http://ollama:11434/api/embeddings',
  body: {
    model: 'nomic-embed-text',
    prompt: question
  },
  json: true
});

const similar = await this.helpers.httpRequest({
  method: 'POST',
  url: 'http://qdrant:6333/collections/docs/points/search',
  body: {
    vector: queryEmbedding.embedding,
    limit: 3
  },
  json: true
});

// 4. Generate answer with context
const context = similar.result.map(r => r.payload.text).join('\n');
const answer = await this.helpers.httpRequest({
  method: 'POST',
  url: 'http://ollama:11434/api/chat',
  body: {
    model: 'llama3.1',
    messages: [
      { role: 'system', content: 'Answer the question based on the context provided.' },
      { role: 'user', content: `Context: ${context}\n\nQuestion: ${question}` }
    ],
    stream: false
  },
  json: true
});

See local RAG for details.

Practical Example 5: Data Extraction

// Extract structured data from text
return {
  prompt: `Extract from the following text: name, email, phone.

Text: ${$input.item.json.text}

Reply as JSON:
{"name": "...", "email": "...", "phone": "..."}`
};

// Ollama with format: json for structured output
{
  "model": "llama3.1",
  "messages": [{"role": "user", "content": "{{$json.prompt}}"}],
  "format": "json",
  "stream": false
}

Error Handling

// Function node with error handling
try {
  const response = await this.helpers.httpRequest({
    method: 'POST',
    url: 'http://ollama:11434/api/chat',
    body: { model: 'llama3.1', messages: [...], stream: false },
    json: true,
    timeout: 30000  // 30 seconds timeout
  });
  return { success: true, result: response.message.content };
} catch (error) {
  if (error.code === 'ECONNREFUSED') {
    return { success: false, error: 'Ollama unreachable' };
  }
  if (error.code === 'ETIMEDOUT') {
    return { success: false, error: 'Timeout' };
  }
  return { success: false, error: error.message };
}

Performance Optimization

1. Model Selection

// For simple tasks: small model
model: 'llama3.1:8b'  // Fast

// For complex tasks: large model
model: 'qwen2.5:32b'  // Slower, but better

2. Caching

// Cache identical requests
const cache = new Map();
const cacheKey = JSON.stringify(prompt);

if (cache.has(cacheKey)) {
  return cache.get(cacheKey);
}

const result = await callOllama(prompt);
cache.set(cacheKey, result);
return result;

3. Batch Processing

// Bundle multiple requests
const batchPrompt = `Classify the following texts:
1. ${texts[0]}
2. ${texts[1]}
3. ${texts[2]}

Reply as JSON array.`;

Security Notes

  • Secure Ollama: Do not expose Ollama to the internet. See API Keys.
  • Validate inputs: Prevent prompt injection. See Prompt Injection.
  • Validate outputs: Check AI responses before processing them.
  • Audit logging: Log all Ollama calls. See Logging.
  • Network isolation: Run n8n and Ollama in a separate Docker network. See Docker Network Isolation.

Common Pitfalls

  • Wrong URL: localhost in Docker refers to the container itself, not the host. Use http://ollama:11434 or host.docker.internal.
  • Timeout too short: Large models need time. Set timeout to 60-120 seconds.
  • No error handling: If Ollama is unreachable, the workflow should not crash.
  • Model not loaded: Check with ollama list whether the model is loaded.
  • Context length exceeded: Long documents must be chunked.
  • No streaming: For long responses, streaming is better, but more complex to handle.

Further Reading

Key Takeaways:

  • n8n sends HTTP requests to the Ollama API.
  • Local AI in workflows without cloud dependencies or API costs.
  • Use cases: classification, summarization, translation, extraction, RAG.
  • Error handling and caching matter.
  • Security: secure Ollama, validate inputs.

FAQ

How do I connect n8n to Ollama?

With an HTTP Request node. URL: http://ollama:11434/api/chat, Method: POST, Body: JSON with model and messages. If both run in Docker, they must be on the same network.

Which model should I use?

For simple tasks: llama3.1:8b (fast). For complex tasks: qwen2.5:32b (better quality). For embeddings: nomic-embed-text. For translation: qwen2.5.

How do I connect n8n and Ollama in Docker?

Create a Docker network (docker network create n8n-ollama), start both containers on the same network, use http://ollama:11434 as the URL in n8n.

Why doesn’t localhost work?

In Docker, localhost refers to the container itself, not the host. Use http://ollama:11434 (container name) or http://host.docker.internal:11434 (host).

How long should the timeout be?

For small models (8B): 30-60 seconds. For large models (32B+): 120-300 seconds. For very long documents: longer.

Can I use streaming?

Yes, but more complex. Set stream: true and process the response incrementally. For simple workflows, stream: false is easier.

What should I do if I encounter errors?

Implement error handling: try/catch, timeout, retry logic. If Ollama is unreachable, the workflow should not crash but execute a fallback action.

How do I build RAG in n8n?

Use Ollama embeddings (nomic-embed-text), store vectors in Qdrant or Pinecone, search for similar documents when questions arrive, generate answers with context.

Is this secure?

Yes, if local. All data stays on your server. But validate inputs for prompt injection and secure Ollama (do not expose to the internet).

What does this cost?

Only hardware costs. n8n and Ollama are open source. You need a machine with enough VRAM for your chosen model. No API costs.

Resources and Further Reading

Back to Blog
Share:

Related Posts