Skip to content
BotServBotServ
DiscordBotOllamaAISlash Commands

Discord Bot with Ollama

Build a Discord bot with Ollama. Slash commands, AI responses, RAG integration, moderation and practical examples.

S

schutzgeist

7 min read
Discord Bot with Ollama

Discord Bot with Ollama

What this article covers

  • How to build a Discord bot with Ollama as your AI backend
  • How Slash Commands, chat responses, and RAG work
  • How to set up and deploy your bot in Discord
  • Practical examples for FAQ bots, moderation, and community support
  • Best practices for rate limiting, security, and moderation

Introduction: Discord bot with Ollama explained

A Discord bot powered by Ollama answers questions in your server using a local AI model, keeping your chat data private and away from OpenAI or other cloud services. Your bot can answer FAQs, search documents via RAG, moderate conversations, and trigger workflows.

This guide is for anyone building a Discord bot with local AI. For background on the tools, check out Ollama and Discord Bot Basics.

Why use a Discord bot with Ollama?

Imagine your community asks the same questions repeatedly: “How do I install X?”, “Where do I find Y?”, “What is Z?”. A bot running Ollama answers these automatically, drawing from your own knowledge base (RAG) and running locally, with no API costs.

How a Discord bot with Ollama works

You create a Discord bot (using discord.py or discord.js) that calls Ollama when messages arrive and posts the response. RAG lets your bot search your documents. Slash Commands give users a structured way to trigger specific features.

The core idea: Discord as the frontend, Ollama as the brain.

Who should read this?

  • Community admins wanting an AI bot
  • Developers building Discord bots
  • Self-hosters running AI locally
  • Teams automating internal support

Basic familiarity with Python or JavaScript, and Ollama, helps.

Key concepts

  • Discord Bot - A program running inside Discord. Use it for: your interface.
  • Ollama - Local model server. Use it for: your AI backend.
  • Slash Commands - Commands like /ask. Use them for: structured invocations.
  • RAG - Knowledge database. Use it for: FAQs.
  • discord.py / discord.js - Bot libraries. Use them for: implementation.
  • Intent - Discord permission. Use it for: message access.
  • Discord Bot Basics - Bot setup. Use it for: getting started.

Setup: Creating your Discord bot

1. Discord Developer Portal

  1. Go to https://discord.com/developers/applications
  2. Click “New Application” and enter a name
  3. Go to the “Bot” tab, click “Add Bot”
  4. Copy your token (keep it secret)
  5. Go to “OAuth2” and select “URL Generator”
    • Scopes: bot, applications.commands
    • Permissions: Send Messages, Read Messages, Use Slash Commands
  6. Open the generated URL and invite the bot to your server

2. Building the bot with discord.py

# bot.py
import discord
from discord import app_commands
import requests
import os

OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://ollama:11434")
TOKEN = os.environ.get("DISCORD_TOKEN")

intents = discord.Intents.default()
intents.message_content = True
client = discord.Client(intents=intents)
tree = app_commands.CommandTree(client)

def call_ollama(prompt, system=None):
    """Call Ollama"""
    messages = []
    if system:
        messages.append({"role": "system", "content": system})
    messages.append({"role": "user", "content": prompt})

    response = requests.post(
        f"{OLLAMA_URL}/api/chat",
        json={
            "model": "llama3.1",
            "messages": messages,
            "stream": False
        },
        timeout=60
    )
    return response.json()["message"]["content"]

@client.event
async def on_ready():
    await tree.sync()
    print(f"Bot online: {client.user}")

# Slash Command: /ask
@tree.command(name="ask", description="Ask the AI bot a question")
async def ask(interaction: discord.Interaction, frage: str):
    await interaction.response.defer()  # Give the AI time to respond

    antwort = call_ollama(
        frage,
        system="You are a helpful assistant on a Discord server. Keep answers brief and useful."
    )

    await interaction.followup.send(f"**Question:** {frage}\n\n{antwort}")

# Reply to messages (optional)
@client.event
async def on_message(message):
    if message.author.bot:
        return

    # Only respond to mentions
    if client.user.mentioned_in(message):
        async with message.channel.typing():
            antwort = call_ollama(
                message.content.replace(f"<@{client.user.id}>", ""),
                system="You are a helpful assistant on Discord."
            )
            await message.reply(antwort)

client.run(TOKEN)

3. Start the bot

pip install discord.py requests

export DISCORD_TOKEN="your-bot-token"
export OLLAMA_URL="http://localhost:11434"

python bot.py

Example 1: FAQ bot with RAG

import chromadb

class FAQBot:
    def __init__(self):
        self.chroma = chromadb.HttpClient(host="chromadb", port=8000)
        self.collection = self.chroma.get_collection("faq")

    async def answer(self, question):
        """Answer a question using RAG"""
        # Find similar documents
        results = self.collection.query(
            query_texts=[question],
            n_results=3
        )

        context = "\n".join(results["documents"][0])

        # Generate answer
        response = call_ollama(
            f"Context:\n{context}\n\nQuestion: {question}",
            system="Answer the question based on the context. Provide sources."
        )

        return {
            "answer": response,
            "sources": results["metadatas"][0]
        }

# Slash Command
@tree.command(name="faq", description="Ask the FAQ")
async def faq(interaction: discord.Interaction, frage: str):
    await interaction.response.defer()
    result = await faq_bot.answer(frage)

    embed = discord.Embed(
        title="FAQ Answer",
        description=result["answer"],
        color=discord.Color.blue()
    )
    embed.set_footer(text=f"Sources: {', '.join(s['source'] for s in result['sources'])}")
    await interaction.followup.send(embed=embed)

Example 2: Moderation bot

@client.event
async def on_message(message):
    if message.author.bot:
        return

    # AI analyzes the message
    analysis = call_ollama(
        f"Analyze this message:\n{message.content}\n\n"
        "Categories: normal, spam, toxic, nsfw, offtopic\n"
        "Reply: CATEGORY|REASON",
        system="You are a moderation assistant."
    )

    category, reason = analysis.split("|", 1)

    if category in ["spam", "toxic", "nsfw"]:
        await message.delete()
        await message.channel.send(
            f"⚠️ Message from {message.author.mention} removed: {reason}"
        )
        # Log
        log_moderation(message, category, reason)

Example 3: Document upload and indexing

@tree.command(name="index", description="Index a document for FAQ")
async def index(interaction: discord.Interaction, anhang: discord.Attachment):
    await interaction.response.defer()

    # Download file
    content = await anhang.read()

    # Extract text
    if anhang.filename.endswith(".pdf"):
        text = extract_pdf(content)
    else:
        text = content.decode("utf-8")

    # Create chunks
    chunks = split_into_chunks(text, 500)

    # Create embeddings
    for i, chunk in enumerate(chunks):
        embedding = get_embedding(chunk)
        collection.add(
            ids=[f"{anhang.filename}_{i}"],
            documents=[chunk],
            embeddings=[embedding],
            metadatas=[{"source": anhang.filename}]
        )

    await interaction.followup.send(
        f"✅ {anhang.filename} indexed ({len(chunks)} chunks)"
    )

Practical Example 4: Triggering a Workflow

@tree.command(name="workflow", description="Workflow auslösen")
async def workflow(interaction: discord.Interaction, name: str, parameter: str):
    await interaction.response.defer()

    # n8n Webhook aufrufen
    response = requests.post(
        f"http://n8n:5678/webhook/{name}",
        json={"parameter": parameter, "user": str(interaction.user)}
    )

    await interaction.followup.send(f"Workflow {name} gestartet: {response.text}")

Docker Setup

version: "3.8"

services:
  discord-bot:
    build: .
    container_name: discord-bot
    restart: unless-stopped
    environment:
      - DISCORD_TOKEN=${DISCORD_TOKEN}
      - OLLAMA_URL=http://ollama:11434
    networks:
      - ai-network
    depends_on:
      - ollama

  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "11434:11434"
    volumes:
      - ollama_data:/root/.ollama
    networks:
      - ai-network

volumes:
  ollama_data:

networks:
  ai-network:
    driver: bridge

Security Considerations

  • Keep bot tokens secret: Never commit them to code or version control. Use environment variables instead. See Secrets.
  • Minimize permissions: Grant the bot only read and write capabilities, never admin rights.
  • Implement rate limiting: Discord enforces rate limits on API calls. Add throttling to your bot.
  • Filter generated content: Language models can produce inappropriate output. Implement content filtering. See Guardrails.
  • Guard against prompt injection: Users can attempt injection attacks. See Prompt Injection.
  • Enable audit logging: Track all bot actions for accountability. See Audit Logging.

Common Pitfalls

  • Hardcoded tokens: Never embed tokens directly in source code. Always use environment variables.
  • Missing intents: Without the message_content intent enabled, your bot won’t be able to read messages.
  • Rate limit collisions: Discord caps API requests globally. Implement retry logic with backoff.
  • Oversized responses: Discord has a 2000-character limit per message. Truncate or split long responses.
  • Slow model responses: Large models take time to generate output. Use defer() for Slash Commands to avoid timeout.
  • Unhandled errors: If Ollama goes down, your bot will crash without proper error handling.

Further Reading

Key Takeaways:

  • Discord bot plus Ollama creates a local AI assistant for your community.
  • Use discord.py for Python or discord.js for JavaScript/TypeScript implementations.
  • Slash Commands provide structured interaction, on_message handles freeform chat.
  • RAG enables document-based Q&A and FAQ systems.
  • Security essentials: protect tokens, enforce rate limiting, filter unsafe content.

FAQ

How do I create a Discord bot?

Register an application in the Discord Developer Portal, add a bot to it, copy the token, and invite the bot to your server. Then implement it using discord.py or discord.js.

Should I use discord.py or discord.js?

discord.py is better for Python, with simpler syntax and solid documentation. discord.js suits Node.js environments and TypeScript projects. Both integrate well with Ollama.

How do I call Ollama from my bot?

Send a POST request to http://ollama:11434/api/chat with your model and messages. Use the requests library in Python or fetch in JavaScript.

Can the bot search documents?

Yes, using RAG. Documents are chunked, embedded, and stored in a vector database. When users ask questions, similar documents are retrieved and provided as context to the model.

Can the bot moderate the server?

Yes. The model analyzes messages for spam, toxicity, and policy violations. When triggered, the bot can delete messages, warn users, or notify admins.

What does this cost?

Discord bots are free. Ollama is open source. You only pay for hardware. No per-message API fees.

Is the bot secure?

Yes, if you keep the token secret, grant minimal permissions, filter outputs, and prevent prompt injection. Running locally means your conversation data stays on your infrastructure.

Are there message limits?

Discord enforces global rate limits of 50 requests per second. Ollama is limited by hardware, typically 2-10 seconds per response. Cache or queue requests at scale.

Sources and Further Resources

Back to Blog
Share:

Related Posts