Telegram Bot with Local AI
What This Article Covers
- Building a Telegram bot with the official Bot API and Ollama.
- BotFather setup, polling vs. webhooks, and token management.
- Complete code examples with aiogram (Python).
- Extensions: inline keyboards, group moderation, RAG, images, and voice messages.
- Docker deployment, systemd service, and production operation.
Introduction
Telegram is the most bot-friendly platform out there. The Bot API is free, officially documented, and requires no business account, approvals, or certificates. You create a bot in two minutes via @BotFather, and it can immediately receive messages, respond, and run its own AI functions.
Combined with Ollama, you get a chat assistant that runs entirely on your server: answer questions, summarize text, search documents with RAG, moderate groups, all without sending chat data to OpenAI or other cloud providers.
Common Use Cases
- Personal AI Assistant: Chat with your local LLM from your phone, on the go, without needing a VPN to a web interface.
- Community Support: Automatically answer recurring questions in your Telegram group.
- Notification Bot: Send monitoring alerts, backup status, or smart home events directly to Telegram.
- Document Assistant: Upload a PDF to the bot, it summarizes the content (RAG + OCR).
- Moderation: Spam detection, welcoming new members, automatic responses to rule violations.
- Agent Interface: Use Telegram as a frontend for your AI agents: trigger tasks via chat.
Prerequisites
- Telegram account
- Server running Ollama (with at least one model loaded, e.g.
llama3.1:8b) - Python 3.10+
- Optional: Docker for deployment, domain + reverse proxy for webhooks
Step 1: Create a Bot with BotFather
Open Telegram and message @BotFather:
/newbot
BotFather asks for a name (display name, e.g. “My AI Assistant”) and a username (must end with bot, e.g. my_ai_bot). You’ll receive a token:
123456789:AAEhBOweik5ad9rQXMENKJBI...
Never hardcode the token in your source. Always use environment variables or a secrets manager. Whoever has the token controls the bot.
Useful BotFather commands for later:
/setdescription # Description in bot profile
/setcommands # Command list (/start, /help, ...)
/setprivacy # Important for groups: disable = bot reads all messages
/setjoingroups # Whether bot can be added to groups
Step 2: Minimal Bot with aiogram
aiogram is the actively maintained Python framework for the Telegram Bot API (async, type hints, router pattern). python-telegram-bot is an alternative; both work fine.
pip install aiogram ollama python-dotenv
.env file:
TELEGRAM_TOKEN=123456789:AAEhBOweik5ad9rQXMENKJBI...
OLLAMA_URL=http://localhost:11434
OLLAMA_MODEL=llama3.1:8b
bot.py:
import asyncio
import os
from dotenv import load_dotenv
from aiogram import Bot, Dispatcher, F
from aiogram.filters import Command
from aiogram.types import Message
import ollama
load_dotenv()
bot = Bot(token=os.environ["TELEGRAM_TOKEN"])
dp = Dispatcher()
ollama_client = ollama.Client(host=os.environ["OLLAMA_URL"])
MODEL = os.environ.get("OLLAMA_MODEL", "llama3.1:8b")
SYSTEM_PROMPT = """You are a helpful Telegram assistant.
Answer briefly and to the point. Maximum 3-4 paragraphs.
Use simple language and explain technical terms."""
@dp.message(Command("start"))
async def cmd_start(message: Message):
await message.answer(
"Hi! I'm an AI bot running a local model. "
"Just send me a message or use /help."
)
@dp.message(Command("help"))
async def cmd_help(message: Message):
await message.answer(
"Commands:\n"
"/start - Start the bot\n"
"/reset - Clear conversation history\n"
"/model - Show current model\n"
"Just type to get AI responses."
)
@dp.message(Command("model"))
async def cmd_model(message: Message):
await message.answer(f"Active model: `{MODEL}`", parse_mode="Markdown")
# Conversation context per user (simple in-memory storage)
conversations: dict[int, list] = {}
@dp.message(Command("reset"))
async def cmd_reset(message: Message):
conversations.pop(message.from_user.id, None)
await message.answer("Conversation cleared.")
@dp.message(F.text)
async def handle_message(message: Message):
user_id = message.from_user.id
# Show "typing..." indicator
await bot.send_chat_action(message.chat.id, "typing")
# Build context (last 10 messages max)
history = conversations.setdefault(user_id, [])
history.append({"role": "user", "content": message.text})
history = history[-10:]
try:
response = ollama_client.chat(
model=MODEL,
messages=[{"role": "system", "content": SYSTEM_PROMPT}] + history,
)
answer = response["message"]["content"]
history.append({"role": "assistant", "content": answer})
conversations[user_id] = history
# Telegram limits messages to 4096 characters
for i in range(0, len(answer), 4000):
await message.answer(answer[i:i + 4000])
except Exception as e:
await message.answer("Error generating response. Is Ollama running?")
print(f"Ollama error: {e}")
async def main():
await dp.start_polling(bot)
if __name__ == "__main__":
asyncio.run(main())
Start the bot:
python bot.py
Send a message to the bot in Telegram, and it will respond with your local model. That’s it: the simplest platform integration you can get.
Polling vs. Webhook
| Mode | How It Works | When to Use |
|---|---|---|
| Polling | Bot asks Telegram for new messages every few seconds | Development, home server without domain, behind NAT |
| Webhook | Telegram pushes updates to your HTTPS endpoint | Production with domain, high message volume, faster responses |
Polling is usually sufficient for self-hosters and much simpler to set up: no reverse proxy, no TLS certificate required. Webhooks make sense with many users or multiple bots on one server.
Webhook setup with aiogram:
from aiogram.webhook.aiohttp_server import SimpleRequestHandler, setup_application
from aiohttp import web
async def main():
await bot.set_webhook(
"https://bots.your-domain.com/telegram/webhook",
secret_token=os.environ["WEBHOOK_SECRET"],
)
app = web.Application()
SimpleRequestHandler(dispatcher=dp, bot=bot).register(app, path="/telegram/webhook")
setup_application(app, dp, bot=bot)
web.run_app(app, host="127.0.0.1", port=8080)
Add a reverse proxy (Nginx/Caddy) with TLS. See Reverse Proxy.
Extensions
Inline Keyboards: Buttons Below Messages
Keyboards work well for choices like “switch model” or “rate response”:
from aiogram.types import InlineKeyboardMarkup, InlineKeyboardButton
from aiogram.filters import Command
@dp.message(Command("modelchoice"))
async def cmd_model_choice(message: Message):
keyboard = InlineKeyboardMarkup(inline_keyboard=[
[InlineKeyboardButton(text="Llama 3.1 (8B)", callback_data="model:llama3.1:8b")],
[InlineKeyboardButton(text="Qwen2.5 (14B)", callback_data="model:qwen2.5:14b")],
[InlineKeyboardButton(text="DeepSeek-R1 (8B)", callback_data="model:deepseek-r1:8b")],
])
await message.answer("Which model?", reply_markup=keyboard)
@dp.callback_query(F.data.startswith("model:"))
async def model_chosen(callback):
model = callback.data.split(":")[1]
user_models[callback.from_user.id] = model
await callback.message.edit_text(f"Model set: `{model}`", parse_mode="Markdown")
await callback.answer()
Group Operation and Moderation
For the bot to read all messages in groups (not just /commands), set /setprivacy to Disable in BotFather. Then:
@dp.message(F.text & F.chat.type.in_({"group", "supergroup"}))
async def group_handler(message: Message):
# Only respond if the bot is mentioned
if f"@{BOT_USERNAME}" in message.text:
await handle_message(message)
return
# Simple AI moderation: classify the message
check = ollama_client.chat(
model=MODEL,
messages=[{
"role": "user",
"content": f"Classify this message as OK or SPAM. "
f"Respond with one word only.\n\n{message.text}"
}],
)
if check["message"]["content"].strip().upper() == "SPAM":
await message.delete()
await message.answer(f"@{message.from_user.username}: Message removed.")
RAG: Bot Answers Questions from Your Documents
The bot searches a vector database before responding. Setup follows the pattern in Local RAG:
from qdrant_client import QdrantClient
qdrant = QdrantClient(host="localhost", port=6333)
def retrieve_context(query: str) -> str:
embedding = ollama_client.embeddings(model="bge-m3", prompt=query)["embedding"]
hits = qdrant.search(collection_name="dokumente", query_vector=embedding, limit=3)
return "\n\n".join(h.payload["text"] for h in hits)
@dp.message(F.text)
async def handle_message(message: Message):
context = retrieve_context(message.text)
response = ollama_client.chat(
model=MODEL,
messages=[
{"role": "system", "content": SYSTEM_PROMPT +
f"\n\nUse this context to answer:\n{context}"},
{"role": "user", "content": message.text},
],
)
await message.answer(response["message"]["content"])
Images and Voice Messages
Telegram delivers files via getFile/download. Send voice messages to Whisper, images to a vision model like llava or minicpm-v:
@dp.message(F.voice)
async def handle_voice(message: Message):
file = await bot.get_file(message.voice.file_id)
path = await bot.download_file(file.file_path)
text = whisper_transcribe(path.name) # local Whisper
await message.answer(f"Got it: „{text}"\n\n")
# then treat as normal text message
@dp.message(F.photo)
async def handle_photo(message: Message):
file = await bot.get_file(message.photo[-1].file_id)
path = await bot.download_file(file.file_path)
response = ollama_client.chat(
model="minicpm-v",
messages=[{
"role": "user",
"content": "Describe this image.",
"images": [path.name],
}],
)
await message.answer(response["message"]["content"])
Streaming Responses
Ollama can stream, and Telegram messages can be edited in place:
sent = await message.answer("…")
text = ""
async for chunk in ollama_client.chat(model=MODEL, messages=msgs, stream=True):
text += chunk["message"]["content"]
if len(text) % 80 == 0: # don't send every token (rate limit!)
await sent.edit_text(text)
await sent.edit_text(text)
Deployment: Docker and systemd
Dockerfile:
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY bot.py .
CMD ["python", "bot.py"]
docker-compose.yml (bot + Ollama + Qdrant in the same stack):
services:
telegram-bot:
build: .
restart: always
env_file: .env
depends_on:
- ollama
networks: [botnet]
ollama:
image: ollama/ollama:latest
restart: always
volumes:
- ollama_data:/root/.ollama
networks: [botnet]
# deploy.resources for GPU see Docker article
volumes:
ollama_data:
networks:
botnet:
Important: In .env, set OLLAMA_URL=http://ollama:11434 - containers reach each other by service name, not localhost.
Security and Privacy
- Token protection: Keep tokens in
.env/Secrets, never in your Git repo. If leaked, use/revokein BotFather. - Access restrictions: Allow only specific user IDs, otherwise anyone can use your bot (and your Ollama):
ALLOWED_USERS = {123456789} # Your Telegram user ID
@dp.message(F.text)
async def handle_message(message: Message):
if message.from_user.id not in ALLOWED_USERS:
return
# ...
- Privacy notice: Messages go to Telegram’s servers (client-server encrypted, not E2E). For truly sensitive conversations, Matrix bots or Signal bots are better.
- Rate limiting: Telegram allows roughly 30 messages per second. For streaming edits, add throttling.
- Cost control: Every request consumes CPU on your server. For public bots, set per-user limits. See Cost control.
Common Pitfalls
- Bot doesn’t respond in groups: Set
/setprivacyto Disable and re-add the bot to the group. - 409 Conflict on startup: A second instance is polling the same token. Kill old processes or containers.
- Message too long: Telegram limit is 4096 characters. Split responses (see code above).
- Markdown errors:
parse_mode="Markdown"throws errors on unbalanced characters from the LLM. Either useparse_mode=Noneor catch the exception. - Memory grows unbounded: Never limit in-memory context per user → long-running bots bloat RAM. Set a limit (see code) or offload to Redis/database.
Further Reading
- IRC-Coding.de: In-depth programming tutorials: bot architectures, Python async, event handling, and more bot projects.
- Chatbot Basics: Rule-based vs. AI, agents vs. chatbots.
- Discord Bot Basics: Alternative platform.
- Ollama API: HTTP interface in detail.
- Local RAG: Document knowledge for your bot.
- Docker: Container fundamentals.
Key takeaways:
- Telegram is the simplest bot platform: BotFather → token → aiogram → done.
- Polling works fine for self-hosters; webhooks only needed at scale or with multiple bots.
- Extensions: inline keyboards, group moderation, RAG, vision, Whisper for voice.
- Restrict access to your own user IDs, treat tokens as secrets.
- Not E2E-encrypted; for sensitive chats, prefer Matrix or Signal.
FAQ
Does a Telegram bot cost anything?
Polling or webhooks?
Why doesn’t the bot respond in groups?
/setprivacy to Disable in BotFather and re-add the bot to the group.aiogram or python-telegram-bot?
Can multiple users use the bot at the same time?
How does the bot remember context?
Are chats private?
Can the bot post to channels?
Resources and Further Reading
- Telegram Bot API: Official reference.
- aiogram: Python framework documentation.
- Ollama: Local model server.
- IRC-Coding.de: Programming tutorials for bot development.


