Running AI Agents with Docker
What this article covers
- How to run AI agents with Docker.
- Why Docker matters for agent deployments.
- Docker Compose for complete agent stacks.
- Practical examples with Ollama, agents, and tools in containers.
- Best practices for isolation, scaling, and maintenance.
Introduction: Understanding Docker for Agents
Docker containerizes your agent stack: Ollama in one container, agent code in another, database in a third. Everything isolated, reproducible, portable. For agents, this means consistent environments, straightforward deployment, and clean isolation between components.
This article is for anyone looking to run agents with Docker. For foundational concepts, see Docker and Running AI Agents Locally.
Why do I need Docker for agents?
Imagine your agent requires Ollama, a database, a vector store, and agent code. Without Docker: install everything manually, configure each piece, hope it works. With Docker: docker-compose up and everything runs, isolated and reproducible.
Docker for agents explained briefly
Docker = a container for each component (Ollama, agent, database, vector store). Docker Compose orchestrates all containers together. Isolated, reproducible, portable.
The core idea: each component runs in its own container, all working as a unified stack.
Who is this article for?
- DevOps engineers deploying agents.
- Self-hosters running agents in containers.
- Developers wanting reproducible environments.
- Teams standardizing agent stacks.
Key concepts
- Docker - Container platform. Use when: isolation is needed.
- Docker Compose - Multi-container orchestration. Use when: managing stacks.
- Ollama - Model server. Use when: running in containers.
- Kubernetes - Container orchestration. Use when: scaling large deployments.
Docker Compose for an agent stack
version: "3.8"
services:
# Model server
ollama:
image: ollama/ollama:latest
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
networks:
- agent_net
# Vector database
qdrant:
image: qdrant/qdrant:latest
ports:
- "6333:6333"
volumes:
- qdrant_data:/qdrant/storage
networks:
- agent_net
# Agent
agent:
build: ./agent
ports:
- "8000:8000"
environment:
- OLLAMA_URL=http://ollama:11434
- QDRANT_URL=http://qdrant:6333
depends_on:
- ollama
- qdrant
networks:
- agent_net
# Workflow tool (optional)
n8n:
image: n8nio/n8n:latest
ports:
- "5678:5678"
volumes:
- n8n_data:/home/node/.n8n
networks:
- agent_net
volumes:
ollama_data:
qdrant_data:
n8n_data:
networks:
agent_net:
driver: bridge
Agent Dockerfile
FROM python:3.11-slim
WORKDIR /app
# Dependencies
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Agent code
COPY . .
# Health check
HEALTHCHECK --interval=30s --timeout=10s --retries=3 \
CMD python -c "import requests; requests.get('http://localhost:8000/health')"
# Start
CMD ["python", "agent.py"]
Practical example: agent in Docker
# agent.py
import os
import requests
from fastapi import FastAPI
app = FastAPI()
OLLAMA_URL = os.getenv("OLLAMA_URL", "http://ollama:11434")
QDRANT_URL = os.getenv("QDRANT_URL", "http://qdrant:6333")
@app.get("/health")
def health():
return {"status": "ok"}
@app.post("/agent")
async def run_agent(task: dict):
"""Agent endpoint"""
# Agent logic here
response = requests.post(f"{OLLAMA_URL}/api/chat", json={
"model": "llama3.1",
"messages": [{"role": "user", "content": task["task"]}],
"tools": tools,
"stream": False
})
return response.json()
Multi-container communication
# Agent communicates with Ollama over the Docker network
OLLAMA_URL = "http://ollama:11434" # Service name as hostname
# Not localhost! Docker networking uses service names
# ollama = container name
# 11434 = port inside the container
Security considerations
- Container isolation: Each container is isolated. A compromised container doesn’t affect others.
- Networking: Keep internal communication on the Docker network, don’t expose it unnecessarily.
- Secrets: Never include secrets in the image. Use environment variables or a secrets management system.
- Resource limits: Set CPU and memory limits so one container can’t consume all resources.
- Updates: Pull new images regularly and rebuild containers with the latest versions.
Common pitfalls
- Localhost vs. service names: Containers communicate via service names, not localhost.
- GPU in containers: For GPU support, use nvidia-docker or —gpus all. Not straightforward.
- Data persistence: Container data is lost on restart. Use volumes for persistent data.
- Networking: Containers need a shared network to communicate.
- Logs: Container logs can grow large. Configure log rotation.
Further reading
- Docker - Docker fundamentals.
- Docker Compose - Multi-container orchestration.
- Ollama Docker - Ollama in Docker.
- Running AI Agents Locally - Overview.
- Kubernetes - For large-scale deployments.
Key Takeaways:
- Docker = containers for each agent component.
- Docker Compose orchestrates Ollama, agent, database, and vector store.
- Isolated, reproducible, portable.
- Use service names for container-to-container communication.
- Docker is the standard for production deployments.
FAQ
Why use Docker for agents?
What do I need for Docker agents?
How do I use GPU with Docker?
How do containers communicate?
Will my data persist after a restart?
Can I scale?
Docker or Kubernetes?
Is Docker secure?
Sources and further reading
- Docker - Official website.
- Docker Compose - Documentation.
- Ollama Docker - Ollama in Docker.


