Skip to content
BotServBotServ
TelegramBotAI BotOllamaaiogramPythonSelf-Hosting

Telegram Bot with Local AI

Build Telegram bots with Ollama: BotFather, polling vs. webhooks, inline keyboards, groups, RAG, Docker deployment.

S

schutzgeist

8 min read
Telegram Bot with Local AI

Telegram Bot with Local AI

What This Article Covers

  • Building a Telegram bot with the official Bot API and Ollama.
  • BotFather setup, polling vs. webhooks, and token management.
  • Complete code examples with aiogram (Python).
  • Extensions: inline keyboards, group moderation, RAG, images, and voice messages.
  • Docker deployment, systemd service, and production operation.

Introduction

Telegram is the most bot-friendly platform out there. The Bot API is free, officially documented, and requires no business account, approvals, or certificates. You create a bot in two minutes via @BotFather, and it can immediately receive messages, respond, and run its own AI functions.

Combined with Ollama, you get a chat assistant that runs entirely on your server: answer questions, summarize text, search documents with RAG, moderate groups, all without sending chat data to OpenAI or other cloud providers.

Common Use Cases

  • Personal AI Assistant: Chat with your local LLM from your phone, on the go, without needing a VPN to a web interface.
  • Community Support: Automatically answer recurring questions in your Telegram group.
  • Notification Bot: Send monitoring alerts, backup status, or smart home events directly to Telegram.
  • Document Assistant: Upload a PDF to the bot, it summarizes the content (RAG + OCR).
  • Moderation: Spam detection, welcoming new members, automatic responses to rule violations.
  • Agent Interface: Use Telegram as a frontend for your AI agents: trigger tasks via chat.

Prerequisites

  • Telegram account
  • Server running Ollama (with at least one model loaded, e.g. llama3.1:8b)
  • Python 3.10+
  • Optional: Docker for deployment, domain + reverse proxy for webhooks

Step 1: Create a Bot with BotFather

Open Telegram and message @BotFather:

/newbot

BotFather asks for a name (display name, e.g. “My AI Assistant”) and a username (must end with bot, e.g. my_ai_bot). You’ll receive a token:

123456789:AAEhBOweik5ad9rQXMENKJBI...

Never hardcode the token in your source. Always use environment variables or a secrets manager. Whoever has the token controls the bot.

Useful BotFather commands for later:

/setdescription     # Description in bot profile
/setcommands        # Command list (/start, /help, ...)
/setprivacy         # Important for groups: disable = bot reads all messages
/setjoingroups      # Whether bot can be added to groups

Step 2: Minimal Bot with aiogram

aiogram is the actively maintained Python framework for the Telegram Bot API (async, type hints, router pattern). python-telegram-bot is an alternative; both work fine.

pip install aiogram ollama python-dotenv

.env file:

TELEGRAM_TOKEN=123456789:AAEhBOweik5ad9rQXMENKJBI...
OLLAMA_URL=http://localhost:11434
OLLAMA_MODEL=llama3.1:8b

bot.py:

import asyncio
import os
from dotenv import load_dotenv
from aiogram import Bot, Dispatcher, F
from aiogram.filters import Command
from aiogram.types import Message
import ollama

load_dotenv()

bot = Bot(token=os.environ["TELEGRAM_TOKEN"])
dp = Dispatcher()
ollama_client = ollama.Client(host=os.environ["OLLAMA_URL"])
MODEL = os.environ.get("OLLAMA_MODEL", "llama3.1:8b")

SYSTEM_PROMPT = """You are a helpful Telegram assistant.
Answer briefly and to the point. Maximum 3-4 paragraphs.
Use simple language and explain technical terms."""


@dp.message(Command("start"))
async def cmd_start(message: Message):
    await message.answer(
        "Hi! I'm an AI bot running a local model. "
        "Just send me a message or use /help."
    )


@dp.message(Command("help"))
async def cmd_help(message: Message):
    await message.answer(
        "Commands:\n"
        "/start - Start the bot\n"
        "/reset - Clear conversation history\n"
        "/model - Show current model\n"
        "Just type to get AI responses."
    )


@dp.message(Command("model"))
async def cmd_model(message: Message):
    await message.answer(f"Active model: `{MODEL}`", parse_mode="Markdown")


# Conversation context per user (simple in-memory storage)
conversations: dict[int, list] = {}

@dp.message(Command("reset"))
async def cmd_reset(message: Message):
    conversations.pop(message.from_user.id, None)
    await message.answer("Conversation cleared.")


@dp.message(F.text)
async def handle_message(message: Message):
    user_id = message.from_user.id

    # Show "typing..." indicator
    await bot.send_chat_action(message.chat.id, "typing")

    # Build context (last 10 messages max)
    history = conversations.setdefault(user_id, [])
    history.append({"role": "user", "content": message.text})
    history = history[-10:]

    try:
        response = ollama_client.chat(
            model=MODEL,
            messages=[{"role": "system", "content": SYSTEM_PROMPT}] + history,
        )
        answer = response["message"]["content"]
        history.append({"role": "assistant", "content": answer})
        conversations[user_id] = history

        # Telegram limits messages to 4096 characters
        for i in range(0, len(answer), 4000):
            await message.answer(answer[i:i + 4000])

    except Exception as e:
        await message.answer("Error generating response. Is Ollama running?")
        print(f"Ollama error: {e}")


async def main():
    await dp.start_polling(bot)

if __name__ == "__main__":
    asyncio.run(main())

Start the bot:

python bot.py

Send a message to the bot in Telegram, and it will respond with your local model. That’s it: the simplest platform integration you can get.

Polling vs. Webhook

ModeHow It WorksWhen to Use
PollingBot asks Telegram for new messages every few secondsDevelopment, home server without domain, behind NAT
WebhookTelegram pushes updates to your HTTPS endpointProduction with domain, high message volume, faster responses

Polling is usually sufficient for self-hosters and much simpler to set up: no reverse proxy, no TLS certificate required. Webhooks make sense with many users or multiple bots on one server.

Webhook setup with aiogram:

from aiogram.webhook.aiohttp_server import SimpleRequestHandler, setup_application
from aiohttp import web

async def main():
    await bot.set_webhook(
        "https://bots.your-domain.com/telegram/webhook",
        secret_token=os.environ["WEBHOOK_SECRET"],
    )
    app = web.Application()
    SimpleRequestHandler(dispatcher=dp, bot=bot).register(app, path="/telegram/webhook")
    setup_application(app, dp, bot=bot)
    web.run_app(app, host="127.0.0.1", port=8080)

Add a reverse proxy (Nginx/Caddy) with TLS. See Reverse Proxy.

Extensions

Inline Keyboards: Buttons Below Messages

Keyboards work well for choices like “switch model” or “rate response”:

from aiogram.types import InlineKeyboardMarkup, InlineKeyboardButton
from aiogram.filters import Command

@dp.message(Command("modelchoice"))
async def cmd_model_choice(message: Message):
    keyboard = InlineKeyboardMarkup(inline_keyboard=[
        [InlineKeyboardButton(text="Llama 3.1 (8B)", callback_data="model:llama3.1:8b")],
        [InlineKeyboardButton(text="Qwen2.5 (14B)", callback_data="model:qwen2.5:14b")],
        [InlineKeyboardButton(text="DeepSeek-R1 (8B)", callback_data="model:deepseek-r1:8b")],
    ])
    await message.answer("Which model?", reply_markup=keyboard)


@dp.callback_query(F.data.startswith("model:"))
async def model_chosen(callback):
    model = callback.data.split(":")[1]
    user_models[callback.from_user.id] = model
    await callback.message.edit_text(f"Model set: `{model}`", parse_mode="Markdown")
    await callback.answer()

Group Operation and Moderation

For the bot to read all messages in groups (not just /commands), set /setprivacy to Disable in BotFather. Then:

@dp.message(F.text & F.chat.type.in_({"group", "supergroup"}))
async def group_handler(message: Message):
    # Only respond if the bot is mentioned
    if f"@{BOT_USERNAME}" in message.text:
        await handle_message(message)
        return

    # Simple AI moderation: classify the message
    check = ollama_client.chat(
        model=MODEL,
        messages=[{
            "role": "user",
            "content": f"Classify this message as OK or SPAM. "
                       f"Respond with one word only.\n\n{message.text}"
        }],
    )
    if check["message"]["content"].strip().upper() == "SPAM":
        await message.delete()
        await message.answer(f"@{message.from_user.username}: Message removed.")

RAG: Bot Answers Questions from Your Documents

The bot searches a vector database before responding. Setup follows the pattern in Local RAG:

from qdrant_client import QdrantClient

qdrant = QdrantClient(host="localhost", port=6333)

def retrieve_context(query: str) -> str:
    embedding = ollama_client.embeddings(model="bge-m3", prompt=query)["embedding"]
    hits = qdrant.search(collection_name="dokumente", query_vector=embedding, limit=3)
    return "\n\n".join(h.payload["text"] for h in hits)


@dp.message(F.text)
async def handle_message(message: Message):
    context = retrieve_context(message.text)
    response = ollama_client.chat(
        model=MODEL,
        messages=[
            {"role": "system", "content": SYSTEM_PROMPT +
             f"\n\nUse this context to answer:\n{context}"},
            {"role": "user", "content": message.text},
        ],
    )
    await message.answer(response["message"]["content"])

Images and Voice Messages

Telegram delivers files via getFile/download. Send voice messages to Whisper, images to a vision model like llava or minicpm-v:

@dp.message(F.voice)
async def handle_voice(message: Message):
    file = await bot.get_file(message.voice.file_id)
    path = await bot.download_file(file.file_path)
    text = whisper_transcribe(path.name)   # local Whisper
    await message.answer(f"Got it: „{text}"\n\n")
    # then treat as normal text message


@dp.message(F.photo)
async def handle_photo(message: Message):
    file = await bot.get_file(message.photo[-1].file_id)
    path = await bot.download_file(file.file_path)
    response = ollama_client.chat(
        model="minicpm-v",
        messages=[{
            "role": "user",
            "content": "Describe this image.",
            "images": [path.name],
        }],
    )
    await message.answer(response["message"]["content"])

Streaming Responses

Ollama can stream, and Telegram messages can be edited in place:

sent = await message.answer("…")
text = ""
async for chunk in ollama_client.chat(model=MODEL, messages=msgs, stream=True):
    text += chunk["message"]["content"]
    if len(text) % 80 == 0:          # don't send every token (rate limit!)
        await sent.edit_text(text)
await sent.edit_text(text)

Deployment: Docker and systemd

Dockerfile:

FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY bot.py .
CMD ["python", "bot.py"]

docker-compose.yml (bot + Ollama + Qdrant in the same stack):

services:
  telegram-bot:
    build: .
    restart: always
    env_file: .env
    depends_on:
      - ollama
    networks: [botnet]

  ollama:
    image: ollama/ollama:latest
    restart: always
    volumes:
      - ollama_data:/root/.ollama
    networks: [botnet]
    # deploy.resources for GPU see Docker article

volumes:
  ollama_data:

networks:
  botnet:

Important: In .env, set OLLAMA_URL=http://ollama:11434 - containers reach each other by service name, not localhost.

Security and Privacy

  • Token protection: Keep tokens in .env/Secrets, never in your Git repo. If leaked, use /revoke in BotFather.
  • Access restrictions: Allow only specific user IDs, otherwise anyone can use your bot (and your Ollama):
ALLOWED_USERS = {123456789}  # Your Telegram user ID

@dp.message(F.text)
async def handle_message(message: Message):
    if message.from_user.id not in ALLOWED_USERS:
        return
    # ...
  • Privacy notice: Messages go to Telegram’s servers (client-server encrypted, not E2E). For truly sensitive conversations, Matrix bots or Signal bots are better.
  • Rate limiting: Telegram allows roughly 30 messages per second. For streaming edits, add throttling.
  • Cost control: Every request consumes CPU on your server. For public bots, set per-user limits. See Cost control.

Common Pitfalls

  • Bot doesn’t respond in groups: Set /setprivacy to Disable and re-add the bot to the group.
  • 409 Conflict on startup: A second instance is polling the same token. Kill old processes or containers.
  • Message too long: Telegram limit is 4096 characters. Split responses (see code above).
  • Markdown errors: parse_mode="Markdown" throws errors on unbalanced characters from the LLM. Either use parse_mode=None or catch the exception.
  • Memory grows unbounded: Never limit in-memory context per user → long-running bots bloat RAM. Set a limit (see code) or offload to Redis/database.

Further Reading

Key takeaways:

  • Telegram is the simplest bot platform: BotFather → token → aiogram → done.
  • Polling works fine for self-hosters; webhooks only needed at scale or with multiple bots.
  • Extensions: inline keyboards, group moderation, RAG, vision, Whisper for voice.
  • Restrict access to your own user IDs, treat tokens as secrets.
  • Not E2E-encrypted; for sensitive chats, prefer Matrix or Signal.

FAQ

Does a Telegram bot cost anything?

No. The Bot API is free and unlimited (subject to rate limits). You only pay for your server hardware. The bot doesn’t need Telegram Premium.

Polling or webhooks?

For self-hosters: polling. No certificate, no reverse proxy, works behind NAT. Webhooks only at high load or when running multiple bots on one server.

Why doesn’t the bot respond in groups?

Privacy mode: by default, bots only see /commands and mentions in groups. Set /setprivacy to Disable in BotFather and re-add the bot to the group.

aiogram or python-telegram-bot?

Both are solid. aiogram is fully async with clean router design and active development, making it the better choice for new projects. python-telegram-bot is older with more legacy tutorials.

Can multiple users use the bot at the same time?

Yes, but local GPU becomes the bottleneck. Ollama processes requests sequentially. At many users, switch to vLLM or queue responses. See Multiple agents.

How does the bot remember context?

Telegram provides no history. You store recent messages per user yourself (in the example, an in-memory dict; for production, use Redis or SQLite). Without custom storage, the bot only knows the current message.

Are chats private?

Partially. Your AI backend stays local, but all messages flow through Telegram’s servers (no E2E in standard chats). For maximum privacy, use Matrix or Signal.

Can the bot post to channels?

Yes. Add the bot as an admin to the channel, then it can post automatically (e.g. monitoring reports, news summaries). In channels, it can only read comments, not posts.

Resources and Further Reading

Back to Blog
Share:

Related Posts