Skip to content
BotServBotServ
MattermostBotAI BotOllamaSelf-HostingSlack AlternativeOpen Source

Mattermost Bot with Local AI

Build a Mattermost bot with Ollama: self-hosted Slack alternative, bot accounts, slash commands, WebSocket and Docker.

S

schutzgeist

6 min read
Mattermost Bot with Local AI

Mattermost Bot with Local AI

What This Article Covers

  • Building a bot on your own Mattermost server with Ollama.
  • Creating bot accounts, working with REST APIs and WebSocket events.
  • Slash commands and interactive posts.
  • Extensions: RAG over company knowledge, moderation, alerting.
  • Why Mattermost plus Ollama gives you complete data sovereignty.

Introduction

Mattermost is the self-hosted Slack alternative: open source (Teams Edition free), runs on your own server, and your data stays with you. Combined with Ollama, you get a fully sovereign solution: chat platform and AI both running locally, nothing leaves your infrastructure.

Bot integration is officially supported and well documented. You get bot accounts with their own tokens, REST APIs for everything, WebSocket for real-time events, slash commands, and interactive message buttons.

Common Use Cases

  • Internal knowledge bot: Company knowledge via RAG, completely on-premise, GDPR-friendly.
  • DevSecOps assistant: Log analysis, runbook answers, incident support.
  • Government and regulated industries: AI chat where cloud services aren’t permitted.
  • Alert hub: Monitoring feeds alerts into Mattermost channels, AI summarizes them.
  • Onboarding: New employees ask the bot without interrupting colleagues.

Prerequisites

  • Running Mattermost server (Docker: mattermost/mattermost-team-edition)
  • Admin access to create bot accounts
  • Ollama with a model on the same host or network
  • Python 3.10+

Step 1: Self-Host Mattermost (Quick Setup)

# docker-compose.yml, Mattermost baseline
services:
  db:
    image: postgres:16
    environment:
      POSTGRES_DB: mattermost
      POSTGRES_USER: mmuser
      POSTGRES_PASSWORD: change-this-password
    volumes: [db_data:/var/lib/postgresql/data]
    restart: always

  mattermost:
    image: mattermost/mattermost-team-edition:latest
    depends_on: [db]
    ports: ["8065:8065"]
    environment:
      MM_SQLSETTINGS_DRIVERNAME: postgres
      MM_SQLSETTINGS_DATASOURCE: "postgres://mmuser:change-this-password@db:5432/mattermost?sslmode=disable"
      MM_SERVICESETTINGS_SITEURL: "https://chat.your-domain.de"
    volumes: [mm_data:/mattermost/data]
    restart: always

volumes:
  db_data:
  mm_data:

In the Admin Panel (System Console → Integrations), enable Bot Accounts.

Step 2: Create a Bot Account

In the UI: System Console → Integrations → Bot Accounts → “Add Bot Account” → Username ki-assistant. Or via API/CLI:

# using mmctl in the container
docker exec mattermost mmctl --local bot create ki-assistant --display-name "AI Assistant" --description "Local AI Bot"

You’ll receive a Bot Token (xxx...), similar to an API key. Store it in .env:

MM_URL=http://mattermost:8065
MM_TOKEN=your-bot-token
MM_TEAM=your-team
OLLAMA_URL=http://ollama:11434
OLLAMA_MODEL=llama3.1:8b

Step 3: Bot with WebSocket (Real-Time)

Mattermost has a straightforward WebSocket event API where messages arrive live:

import asyncio
import json
import os
import aiohttp
import websockets
import ollama
from dotenv import load_dotenv

load_dotenv()
MM_URL = os.environ["MM_URL"]
MM_TOKEN = os.environ["MM_TOKEN"]
MODEL = os.environ.get("OLLAMA_MODEL", "llama3.1:8b")

ollama_client = ollama.Client(host=os.environ["OLLAMA_URL"])
HEADERS = {"Authorization": f"Bearer {MM_TOKEN}"}

SYSTEM = "You are the internal AI assistant. Keep responses concise."


async def api(method, path, payload=None):
    async with aiohttp.ClientSession() as s:
        async with s.request(method, f"{MM_URL}/api/v4{path}",
                             headers=HEADERS, json=payload) as r:
            return await r.json()


async def post_message(channel_id, text, root_id=None):
    await api("POST", "/posts", {
        "channel_id": channel_id,
        "message": text[:16000],
        "root_id": root_id,
    })


def ask_ollama(text, history=None):
    msgs = [{"role": "system", "content": SYSTEM}]
    if history:
        msgs += history
    msgs.append({"role": "user", "content": text})
    return ollama_client.chat(model=MODEL, messages=msgs)["message"]["content"]


async def main():
    me = await api("GET", "/users/me")
    bot_id = me["id"]

    ws_url = MM_URL.replace("http", "ws") + "/api/v4/websocket"
    async with websockets.connect(ws_url) as ws:
        await ws.send(json.dumps({
            "seq": 1, "action": "authentication_challenge",
            "data": {"token": MM_TOKEN}
        }))
        async for raw in ws:
            evt = json.loads(raw)
            if evt.get("event") != "posted":
                continue
            post = json.loads(evt["data"]["post"])
            if post["user_id"] == bot_id:   # ignore own posts
                continue

            text = post["message"]
            mention = f"@{me['username']}"
            channel_type = evt["data"].get("channel_type")

            # Only respond to direct messages or @mentions
            if channel_type != "D" and mention not in text:
                continue
            text = text.replace(mention, "").strip()
            if not text:
                continue

            # Load thread context from root posts
            history = []
            if post.get("root_id"):
                thread = await api("GET", f"/posts/{post['root_id']}")
                for p in thread["posts"].values():
                    role = "assistant" if p["user_id"] == bot_id else "user"
                    history.append({"role": role, "content": p["message"]})

            answer = ask_ollama(text, history)
            await post_message(post["channel_id"], answer,
                               root_id=post.get("root_id") or post["id"])

asyncio.run(main())

Slash Commands

Register a slash command /ai in the Admin Console pointing to your HTTP endpoint. Mattermost will push requests to you, so your bot needs an HTTP port:

from aiohttp import web

async def slash_ai(request):
    data = await request.post()
    prompt = data.get("text", "").strip()
    if not prompt:
        return web.json_response({"text": "Usage: `/ai <question>`",
                                  "response_type": "ephemeral"})
    answer = ask_ollama(prompt)
    return web.json_response({
        "text": answer,
        "response_type": "ephemeral",   # only the user who ran the command sees it
    })

app = web.Application()
app.router.add_post("/slash/ai", slash_ai)
web.run_app(app, port=8080)

Extensions

RAG Over Company Knowledge

Identical to the Slack bot approach: embed the question, search your company-knowledge collection in Qdrant, inject results into the system prompt. With Mattermost it’s even easier to justify since all data is already on-premise.

Interactive Posts (Buttons)

await api("POST", "/posts", {
    "channel_id": channel_id,
    "message": "Was this answer helpful?",
    "props": {"attachments": [{
        "actions": [
            {"id": "good", "name": "👍 Helpful", "type": "button",
             "integration": {"url": f"{BOT_URL}/action", "context": {"vote": "up"}}},
            {"id": "bad", "name": "👎 Wrong", "type": "button",
             "integration": {"url": f"{BOT_URL}/action", "context": {"vote": "down"}}},
        ]
    }]},
})

Proactive Alerts

The bot posts directly to channels via monitoring integration:

# Cron/Systemd timer or event handler calls:
async def alert(channel_id, service, status):
    summary = ask_ollama(
        f"Summarize this monitoring alert briefly and suggest the most "
        f"likely fix:\nService: {service}\nStatus: {status}"
    )
    await post_message(channel_id, f"🚨 **{service}**\n{summary}")

Moderation

Classify new posts as OK/SPAM/TOXIC, delete violations using DELETE /posts/{id}, and warn the user. Follow the same pattern as the Telegram article.

Deployment: Everything in One Stack

services:
  mm-bot:
    build: ./bot
    restart: always
    env_file: .env
    depends_on: [mattermost, ollama]
    ports: ["127.0.0.1:8080:8080"]   # for slash commands/actions

  ollama:
    image: ollama/ollama:latest
    volumes: [ollama_data:/root/.ollama]
    restart: always

Bot, Mattermost, and Ollama run in the same Docker network, so there are no external dependencies and no data leaves your infrastructure.

Security

  • Bot Token: Full bot privileges, treat like an admin token. Store in .env or Docker Secrets.
  • Channel Isolation: The bot only sees channels it’s a member of, not everything by default.
  • Rate Limiting: For public teams, enforce per-user access limits using token bucket.
  • Logging: Log all bot invocations (Audit Logging).
  • TLS: Use HTTPS for slash commands and actions; HTTP is fine within the Docker network.

Common Pitfalls

  • Bot doesn’t see a channel: The bot must be added as a channel member, either via UI or POST /channels/{id}/members.
  • WebSocket disconnects: Implement reconnect logic, network blips are expected.
  • Self-reply loops: Check post.user_id == bot_id to prevent the bot from answering itself.
  • Incorrect SiteURL: Interactive buttons and attachments require the correct SiteURL.
  • Message limits: Posts are capped at ~16,383 characters; split long responses.

Further Reading

Key Takeaways:

  • Mattermost + Ollama = a completely sovereign AI chat platform, 100% on-premise.
  • Bot account + REST/WebSocket is all you need; no app approval process like Slack.
  • WebSocket events for real-time updates, REST for posting, slash commands via HTTP endpoint.
  • Ideal for companies, government agencies, and anyone who won’t hand over chat data.
  • Team Edition is free and unlimited.

FAQ

Which Mattermost edition do I need?

The free Team Edition covers everything: bots, slash commands, and WebSocket are all included. Professional/Enterprise is only needed for SAML, compliance exports, and clustering.

Mattermost or Slack?

Choose Mattermost if data sovereignty matters (everything on-premise, no cloud). Choose Slack if your team is already there and you don’t want to manage infrastructure. The API concepts are nearly identical, so code is portable.

Bot account or regular user?

Always use a bot account: it gets its own token, is marked as a bot in the UI, cannot log in with a password, and has clearer permissions. Repurposing normal user accounts as bots violates Mattermost policy.

How many resources does Mattermost need?

Mattermost itself uses ~1 GB RAM and minimal CPU. Postgres requires ~500 MB. Ollama with a model (8B Q4 format) needs ~6 GB VRAM. A small PC with 16 GB RAM and a GPU handles small teams comfortably.

Can the bot see all channels?

No, only channels it’s a member of, plus direct messages. This is by design: access control happens through channel membership, not code.

Bots or Mattermost plugins?

For AI integrations, a bot is sufficient (simpler, language-agnostic, independently deployable). Plugins (in Go/React) are only needed for deep UI integration. There are also official AI plugins available (mattermost-ai).

What do I need to back up?

The Postgres database (chats), /mattermost/data (uploads), and your bot session/config. See the backup article. The bot code itself is stateless.

Sources and Further Reading

Back to Blog
Share:

Related Posts