Discord Bot with Ollama
What this article covers
- How to build a Discord bot with Ollama as your AI backend
- How Slash Commands, chat responses, and RAG work
- How to set up and deploy your bot in Discord
- Practical examples for FAQ bots, moderation, and community support
- Best practices for rate limiting, security, and moderation
Introduction: Discord bot with Ollama explained
A Discord bot powered by Ollama answers questions in your server using a local AI model, keeping your chat data private and away from OpenAI or other cloud services. Your bot can answer FAQs, search documents via RAG, moderate conversations, and trigger workflows.
This guide is for anyone building a Discord bot with local AI. For background on the tools, check out Ollama and Discord Bot Basics.
Why use a Discord bot with Ollama?
Imagine your community asks the same questions repeatedly: “How do I install X?”, “Where do I find Y?”, “What is Z?”. A bot running Ollama answers these automatically, drawing from your own knowledge base (RAG) and running locally, with no API costs.
How a Discord bot with Ollama works
You create a Discord bot (using discord.py or discord.js) that calls Ollama when messages arrive and posts the response. RAG lets your bot search your documents. Slash Commands give users a structured way to trigger specific features.
The core idea: Discord as the frontend, Ollama as the brain.
Who should read this?
- Community admins wanting an AI bot
- Developers building Discord bots
- Self-hosters running AI locally
- Teams automating internal support
Basic familiarity with Python or JavaScript, and Ollama, helps.
Key concepts
- Discord Bot - A program running inside Discord. Use it for: your interface.
- Ollama - Local model server. Use it for: your AI backend.
- Slash Commands - Commands like /ask. Use them for: structured invocations.
- RAG - Knowledge database. Use it for: FAQs.
- discord.py / discord.js - Bot libraries. Use them for: implementation.
- Intent - Discord permission. Use it for: message access.
- Discord Bot Basics - Bot setup. Use it for: getting started.
Setup: Creating your Discord bot
1. Discord Developer Portal
- Go to https://discord.com/developers/applications
- Click “New Application” and enter a name
- Go to the “Bot” tab, click “Add Bot”
- Copy your token (keep it secret)
- Go to “OAuth2” and select “URL Generator”
- Scopes:
bot,applications.commands - Permissions:
Send Messages,Read Messages,Use Slash Commands
- Scopes:
- Open the generated URL and invite the bot to your server
2. Building the bot with discord.py
# bot.py
import discord
from discord import app_commands
import requests
import os
OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://ollama:11434")
TOKEN = os.environ.get("DISCORD_TOKEN")
intents = discord.Intents.default()
intents.message_content = True
client = discord.Client(intents=intents)
tree = app_commands.CommandTree(client)
def call_ollama(prompt, system=None):
"""Call Ollama"""
messages = []
if system:
messages.append({"role": "system", "content": system})
messages.append({"role": "user", "content": prompt})
response = requests.post(
f"{OLLAMA_URL}/api/chat",
json={
"model": "llama3.1",
"messages": messages,
"stream": False
},
timeout=60
)
return response.json()["message"]["content"]
@client.event
async def on_ready():
await tree.sync()
print(f"Bot online: {client.user}")
# Slash Command: /ask
@tree.command(name="ask", description="Ask the AI bot a question")
async def ask(interaction: discord.Interaction, frage: str):
await interaction.response.defer() # Give the AI time to respond
antwort = call_ollama(
frage,
system="You are a helpful assistant on a Discord server. Keep answers brief and useful."
)
await interaction.followup.send(f"**Question:** {frage}\n\n{antwort}")
# Reply to messages (optional)
@client.event
async def on_message(message):
if message.author.bot:
return
# Only respond to mentions
if client.user.mentioned_in(message):
async with message.channel.typing():
antwort = call_ollama(
message.content.replace(f"<@{client.user.id}>", ""),
system="You are a helpful assistant on Discord."
)
await message.reply(antwort)
client.run(TOKEN)
3. Start the bot
pip install discord.py requests
export DISCORD_TOKEN="your-bot-token"
export OLLAMA_URL="http://localhost:11434"
python bot.py
Example 1: FAQ bot with RAG
import chromadb
class FAQBot:
def __init__(self):
self.chroma = chromadb.HttpClient(host="chromadb", port=8000)
self.collection = self.chroma.get_collection("faq")
async def answer(self, question):
"""Answer a question using RAG"""
# Find similar documents
results = self.collection.query(
query_texts=[question],
n_results=3
)
context = "\n".join(results["documents"][0])
# Generate answer
response = call_ollama(
f"Context:\n{context}\n\nQuestion: {question}",
system="Answer the question based on the context. Provide sources."
)
return {
"answer": response,
"sources": results["metadatas"][0]
}
# Slash Command
@tree.command(name="faq", description="Ask the FAQ")
async def faq(interaction: discord.Interaction, frage: str):
await interaction.response.defer()
result = await faq_bot.answer(frage)
embed = discord.Embed(
title="FAQ Answer",
description=result["answer"],
color=discord.Color.blue()
)
embed.set_footer(text=f"Sources: {', '.join(s['source'] for s in result['sources'])}")
await interaction.followup.send(embed=embed)
Example 2: Moderation bot
@client.event
async def on_message(message):
if message.author.bot:
return
# AI analyzes the message
analysis = call_ollama(
f"Analyze this message:\n{message.content}\n\n"
"Categories: normal, spam, toxic, nsfw, offtopic\n"
"Reply: CATEGORY|REASON",
system="You are a moderation assistant."
)
category, reason = analysis.split("|", 1)
if category in ["spam", "toxic", "nsfw"]:
await message.delete()
await message.channel.send(
f"⚠️ Message from {message.author.mention} removed: {reason}"
)
# Log
log_moderation(message, category, reason)
Example 3: Document upload and indexing
@tree.command(name="index", description="Index a document for FAQ")
async def index(interaction: discord.Interaction, anhang: discord.Attachment):
await interaction.response.defer()
# Download file
content = await anhang.read()
# Extract text
if anhang.filename.endswith(".pdf"):
text = extract_pdf(content)
else:
text = content.decode("utf-8")
# Create chunks
chunks = split_into_chunks(text, 500)
# Create embeddings
for i, chunk in enumerate(chunks):
embedding = get_embedding(chunk)
collection.add(
ids=[f"{anhang.filename}_{i}"],
documents=[chunk],
embeddings=[embedding],
metadatas=[{"source": anhang.filename}]
)
await interaction.followup.send(
f"✅ {anhang.filename} indexed ({len(chunks)} chunks)"
)
Practical Example 4: Triggering a Workflow
@tree.command(name="workflow", description="Workflow auslösen")
async def workflow(interaction: discord.Interaction, name: str, parameter: str):
await interaction.response.defer()
# n8n Webhook aufrufen
response = requests.post(
f"http://n8n:5678/webhook/{name}",
json={"parameter": parameter, "user": str(interaction.user)}
)
await interaction.followup.send(f"Workflow {name} gestartet: {response.text}")
Docker Setup
version: "3.8"
services:
discord-bot:
build: .
container_name: discord-bot
restart: unless-stopped
environment:
- DISCORD_TOKEN=${DISCORD_TOKEN}
- OLLAMA_URL=http://ollama:11434
networks:
- ai-network
depends_on:
- ollama
ollama:
image: ollama/ollama:latest
container_name: ollama
restart: unless-stopped
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
networks:
- ai-network
volumes:
ollama_data:
networks:
ai-network:
driver: bridge
Security Considerations
- Keep bot tokens secret: Never commit them to code or version control. Use environment variables instead. See Secrets.
- Minimize permissions: Grant the bot only read and write capabilities, never admin rights.
- Implement rate limiting: Discord enforces rate limits on API calls. Add throttling to your bot.
- Filter generated content: Language models can produce inappropriate output. Implement content filtering. See Guardrails.
- Guard against prompt injection: Users can attempt injection attacks. See Prompt Injection.
- Enable audit logging: Track all bot actions for accountability. See Audit Logging.
Common Pitfalls
- Hardcoded tokens: Never embed tokens directly in source code. Always use environment variables.
- Missing intents: Without the
message_contentintent enabled, your bot won’t be able to read messages. - Rate limit collisions: Discord caps API requests globally. Implement retry logic with backoff.
- Oversized responses: Discord has a 2000-character limit per message. Truncate or split long responses.
- Slow model responses: Large models take time to generate output. Use
defer()for Slash Commands to avoid timeout. - Unhandled errors: If Ollama goes down, your bot will crash without proper error handling.
Further Reading
- Discord Bot Basics - Core concepts for building bots.
- Ollama - Local model server.
- Local RAG - Knowledge base integration.
- Discord.py - Python bot library.
- Prompt Injection - Security best practices.
- Guardrails - Content filtering.
Key Takeaways:
- Discord bot plus Ollama creates a local AI assistant for your community.
- Use discord.py for Python or discord.js for JavaScript/TypeScript implementations.
- Slash Commands provide structured interaction, on_message handles freeform chat.
- RAG enables document-based Q&A and FAQ systems.
- Security essentials: protect tokens, enforce rate limiting, filter unsafe content.
FAQ
How do I create a Discord bot?
Should I use discord.py or discord.js?
How do I call Ollama from my bot?
Can the bot search documents?
Can the bot moderate the server?
What does this cost?
Is the bot secure?
Are there message limits?
Sources and Further Resources
- Discord Developer Portal - Create and manage bots.
- discord.py - Python library documentation.
- discord.js - JavaScript library documentation.
- Ollama - Local model server.


