Skip to content
BotServBotServ
SlackBotAI BotOllamaBoltSocket ModeTeam AI

Slack Bot with Local AI

Build a Slack bot with Ollama: Bolt framework, Socket Mode, slash commands, threads, and extensions.

S

schutzgeist

7 min read
Slack Bot with Local AI

Slack Bot with Local AI

What This Article Covers

  • Building a Slack app with the Bolt framework and Ollama
  • Socket Mode vs. HTTP Events: why Socket Mode is ideal for self-hosters
  • Slash commands, mentions, thread context, and interactive blocks
  • Extensions: RAG, file uploads, multi-user context
  • Deployment and security in enterprise environments

Introduction

Slack is the standard in tech teams and offers a solid official bot API. Slack apps support slash commands, mentions, buttons, modals, and threads. The key advantage for self-hosters: Socket Mode connects your bot to Slack via WebSocket, eliminating the need for a public endpoint, reverse proxy, or firewall holes.

With Ollama backing it, your Slack bot becomes an internal team assistant: company knowledge via RAG, code questions, summaries, all without sending Slack content to AI clouds (though Slack itself still sees the messages, of course).

Common Use Cases

  • Internal knowledge bot: “How do I request time off?” → Bot responds from the wiki via RAG
  • DevOps assistant: “/ai why is the build failing?” → Analyze log snippets
  • Onboarding helper: New employees ask the bot instead of colleagues
  • Code review bot: Upload a diff, get review comments from a local coding model
  • Meeting summaries: Condense thread discussions automatically
  • IT helpdesk first line: Catch frequent questions, escalate the rest

Prerequisites

  • Slack workspace (free tier works, but note: message history limited to 90 days on free plans)
  • Admin rights in the workspace (or permission to install apps)
  • Server running Ollama + Python 3.10+

Step 1: Create a Slack App

On api.slack.com/apps → “Create New App” → “From scratch”:

  1. Enable Socket Mode (Settings → Socket Mode → Enable) → Generate app-level token xapp-... (scope: connections:write)
  2. Bot Token Scopes (OAuth & Permissions): app_mentions:read, chat:write, commands, channels:history, im:history, files:read
  3. Create Slash Command: /ai (request URL doesn’t matter for Socket Mode; use a placeholder)
  4. Event Subscriptions: Subscribe to app_mention, message.im
  5. Install to Workspace → Copy bot token xoxb-...

You’ll have two tokens: xapp- (app-level, for Socket Mode) and xoxb- (bot token). Add both to .env:

SLACK_BOT_TOKEN=xoxb-...
SLACK_APP_TOKEN=xapp-...
OLLAMA_URL=http://localhost:11434
OLLAMA_MODEL=llama3.1:8b

Step 2: Bot with Bolt (Python)

pip install slack-bolt ollama python-dotenv

bot.py:

import os
import ollama
from dotenv import load_dotenv
from slack_bolt import App
from slack_bolt.adapter.socket_mode import SocketModeHandler

load_dotenv()

app = App(token=os.environ["SLACK_BOT_TOKEN"])
ollama_client = ollama.Client(host=os.environ["OLLAMA_URL"])
MODEL = os.environ.get("OLLAMA_MODEL", "llama3.1:8b")

SYSTEM = ("You are the internal AI assistant. Answer briefly and precisely. "
          "For code, use Markdown code blocks.")


def ask_ollama(prompt: str, history: list | None = None) -> str:
    messages = [{"role": "system", "content": SYSTEM}]
    if history:
        messages.extend(history)
    messages.append({"role": "user", "content": prompt})
    resp = ollama_client.chat(model=MODEL, messages=messages)
    return resp["message"]["content"]


# Load thread history, Slack provides it directly!
def thread_history(client, channel, thread_ts) -> list:
    if not thread_ts:
        return []
    result = client.conversations_replies(channel=channel, ts=thread_ts, limit=20)
    history = []
    for m in result["messages"][1:]:   # first message = current message
        role = "assistant" if m.get("bot_id") else "user"
        history.append({"role": role, "content": m.get("text", "")})
    return history


@app.event("app_mention")
def handle_mention(event, say, client):
    history = thread_history(client, event["channel"], event.get("thread_ts"))
    # Strip mention markup <@U123>
    text = event["text"].split(">", 1)[-1].strip()
    answer = ask_ollama(text, history)
    say(text=answer, thread_ts=event.get("thread_ts") or event["ts"])


@app.event("message")
def handle_dm(event, say, client):
    # Only direct messages (channel type "im"), skip bot's own messages
    if event.get("channel_type") != "im" or event.get("bot_id"):
        return
    history = thread_history(client, event["channel"], event.get("thread_ts"))
    answer = ask_ollama(event.get("text", ""), history)
    say(text=answer, thread_ts=event.get("thread_ts") or event["ts"])


@app.command("/ai")
def cmd_ai(ack, respond, command):
    ack()   # Slack requires acknowledgment within 3 seconds
    prompt = command.get("text", "").strip()
    if not prompt:
        respond("Usage: `/ai <question>`")
        return
    # Response may take longer than 3s, so use delayed respond
    answer = ask_ollama(prompt)
    respond(answer)   # Ephemeral by default: only the user who asked sees it


if __name__ == "__main__":
    SocketModeHandler(app, os.environ["SLACK_APP_TOKEN"]).start()

That’s it: no web server, no port, no domain. The bot connects to Slack via WebSocket and is immediately available.

Socket Mode vs. HTTP Events

AspectSocket ModeHTTP Events API
ConnectionBot establishes WebSocket to SlackSlack pushes to your HTTPS endpoint
Public EndpointNot requiredRequired (domain + TLS + reverse proxy)
Firewall/NATWorks everywhereRequires inbound ports
ScalingOne connection per processUnlimited instances
Best ForSelf-hosting, internal botsPublic Slack Marketplace apps

For internal team bots, Socket Mode is clearly the better choice.

Extensions

Interactive Blocks (Buttons, Modals)

@app.command("/ai")
def cmd_ai(ack, respond, command):
    ack()
    respond(blocks=[
        {"type": "section", "text": {"type": "mrkdwn",
          "text": f"*Question:* {command['text']}"}},
        {"type": "actions", "elements": [
            {"type": "button", "text": {"type": "plain_text",
             "text": "As Code"}, "action_id": "fmt_code",
             "value": command["text"]},
            {"type": "button", "text": {"type": "plain_text",
             "text": "Detailed"}, "action_id": "fmt_long",
             "value": command["text"]},
        ]},
    ])

@app.action("fmt_code")
def on_code(ack, body, respond):
    ack()
    prompt = body["actions"][0]["value"]
    answer = ask_ollama(prompt + "\nRespond with only a code block.")
    respond(answer)

RAG: Internal Company Knowledge

The bot first searches the vector database (wiki exports, handbooks, runbooks chunked with bge-m3):

def ask_with_rag(prompt, history=None):
    emb = ollama_client.embeddings(model="bge-m3", prompt=prompt)["embedding"]
    hits = qdrant.search("company_knowledge", query_vector=emb, limit=4)
    context = "\n\n".join(h.payload["text"] for h in hits)
    system = SYSTEM + f"\n\nRelevant internal documents:\n{context}\n" \
              "Cite the source at the end of your answer."
    # ...

Important: Depending on channel or DM, access controls may be necessary. Not every channel should see every document. Build access filters into the search query (filter={"must": [{"key": "allowed_channels"...}]}).

Analyzing File Uploads

A user uploads a file → bot retrieves it from Slack → passes it to the AI:

@app.event("message")
def handle_file(event, say, client):
    for f in event.get("files", []):
        if f["filetype"] in ("png", "jpg"):
            # Vision model
            data = download_slack_file(f["url_private"])
            resp = ollama_client.chat(model="minicpm-v", messages=[{
                "role": "user", "content": "Was ist auf diesem Bild?",
                "images": [data],
            }])
            say(resp["message"]["content"], thread_ts=event["ts"])

Streaming Responses: Message Updates

Slack’s chat.update method lets you rewrite responses incrementally into the same message:

posted = say("…", thread_ts=event["ts"])
text = ""
for chunk in ollama_client.chat(model=MODEL, messages=msgs, stream=True):
    text += chunk["message"]["content"]
    if len(text) % 100 == 0:
        client.chat_update(channel=channel, ts=posted["ts"], text=text)
client.chat_update(channel=channel, ts=posted["ts"], text=text)

Deployment

FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY bot.py .
CMD ["python", "bot.py"]
services:
  slack-bot:
    build: .
    restart: always
    env_file: .env
    depends_on: [ollama]
  ollama:
    image: ollama/ollama:latest
    volumes: [ollama_data:/root/.ollama]
volumes:
  ollama_data:

Socket Mode requires only outgoing connections, so your bot runs behind any firewall, even in a home LAN without port forwarding.

Security and Compliance

  • Two tokens, both secret: xoxb (bot permissions) and xapp (Socket connection). Revoke both immediately if exposed on api.slack.com.
  • Minimize scopes: Request only what the bot needs. Include channels:history only if you need thread context in channels.
  • GDPR: Slack is a US service (Salesforce), so chat content goes to Slack servers. For confidential internal communication, self-host Mattermost or Matrix.
  • Audit logging: Log all bot queries (who asked what, which documents were retrieved). See Audit Logging.
  • Channel isolation: Don’t let your bot expose every internal document to every channel. Mirror access controls in your RAG filter.

Common Pitfalls

  • 3-second ACK deadline: Slash commands and interactive actions must acknowledge in under 3 seconds, otherwise you’ll get an operation_timeout error. Always call ack() first, then send the slow response via respond().
  • Mention markup: event["text"] contains raw mentions like <@U0123ABC>. Strip these before passing text to your AI prompt.
  • Message events flood: The message event fires for every visible message in channels. Filter by channel_type=="im" for DMs, mention detection, or a prefix pattern like !ai, otherwise your bot processes every single channel post.
  • Duplicate events: Slack sometimes resends events on retry. Implement deduplication using event_id.
  • Free plan limits: Only 90 days of history. Older thread messages disappear, breaking thread context recovery.

Further Reading

Key Takeaways:

  • Socket Mode = no public endpoint required, ideal for self-hosters.
  • Bolt framework (Python/JavaScript) is the official, simplest approach.
  • Slash commands need ACK within 3 seconds; deliver the response via respond() afterward.
  • Slack provides thread history automatically, so context comes nearly free.
  • For sensitive data, self-host Mattermost or Matrix instead of using Slack.

FAQ

What is Socket Mode?

Your bot initiates an outgoing WebSocket connection to Slack, and Slack pushes events through it. You don’t need a public HTTPS endpoint, which makes it perfect for behind-firewall and NAT setups without port forwarding.

Do I need admin rights?

To install the app, yes (or a workspace admin must approve it). Many corporate workspaces restrict app installation, so check with your workspace admin beforehand.

Can only the person asking see the response?

Yes. Slash command responses are “ephemeral” by default (visible only to the user who triggered them). To show the response to everyone, set response_type: in_channel or use say().

How does the bot get thread context?

Slack provides it for you: conversations_replies(channel, thread_ts) returns all messages in a thread. You don’t need to store anything yourself, a major advantage over Telegram or Signal.

Is a Slack bot GDPR compliant?

Slack stores messages on US servers (EU hosting available on Enterprise plans). It’s standard and reasonable for internal team bots, but review it for customer data. For maximum data sovereignty, self-host Mattermost.

What does it cost?

The Slack API is free, even on the free plan. Only your server costs apply. The free plan does cap history at 90 days though, so for RAG over documentation, use your own data source instead.

Multiple concurrent users?

Yes, Bolt is async-capable. The bottleneck is Ollama, which runs inference serially. With many users, switch to a vLLM backend or add a queue with a wait message.

The bot responds to everything. How do I filter?

The message event fires for every message in visible channels. Filter on channel_type=='im' for DMs, on mentions (@bot) in channels, or on a prefix like !ai.

Resources and Further Reading

Back to Blog
Share:

Related Posts