Skip to content
BotServBotServ
DockerAI AgentContainerDocker-ComposeOrchestration

Running AI Agents with Docker

Run AI agents with Docker. Containerization, Docker-Compose, orchestration and practical examples.

S

schutzgeist

4 min read
Running AI Agents with Docker

Running AI Agents with Docker

What this article covers

  • How to run AI agents with Docker.
  • Why Docker matters for agent deployments.
  • Docker Compose for complete agent stacks.
  • Practical examples with Ollama, agents, and tools in containers.
  • Best practices for isolation, scaling, and maintenance.

Introduction: Understanding Docker for Agents

Docker containerizes your agent stack: Ollama in one container, agent code in another, database in a third. Everything isolated, reproducible, portable. For agents, this means consistent environments, straightforward deployment, and clean isolation between components.

This article is for anyone looking to run agents with Docker. For foundational concepts, see Docker and Running AI Agents Locally.

Why do I need Docker for agents?

Imagine your agent requires Ollama, a database, a vector store, and agent code. Without Docker: install everything manually, configure each piece, hope it works. With Docker: docker-compose up and everything runs, isolated and reproducible.

Docker for agents explained briefly

Docker = a container for each component (Ollama, agent, database, vector store). Docker Compose orchestrates all containers together. Isolated, reproducible, portable.

The core idea: each component runs in its own container, all working as a unified stack.

Who is this article for?

  • DevOps engineers deploying agents.
  • Self-hosters running agents in containers.
  • Developers wanting reproducible environments.
  • Teams standardizing agent stacks.

Key concepts

  • Docker - Container platform. Use when: isolation is needed.
  • Docker Compose - Multi-container orchestration. Use when: managing stacks.
  • Ollama - Model server. Use when: running in containers.
  • Kubernetes - Container orchestration. Use when: scaling large deployments.

Docker Compose for an agent stack

version: "3.8"

services:
  # Model server
  ollama:
    image: ollama/ollama:latest
    ports:
      - "11434:11434"
    volumes:
      - ollama_data:/root/.ollama
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
    networks:
      - agent_net

  # Vector database
  qdrant:
    image: qdrant/qdrant:latest
    ports:
      - "6333:6333"
    volumes:
      - qdrant_data:/qdrant/storage
    networks:
      - agent_net

  # Agent
  agent:
    build: ./agent
    ports:
      - "8000:8000"
    environment:
      - OLLAMA_URL=http://ollama:11434
      - QDRANT_URL=http://qdrant:6333
    depends_on:
      - ollama
      - qdrant
    networks:
      - agent_net

  # Workflow tool (optional)
  n8n:
    image: n8nio/n8n:latest
    ports:
      - "5678:5678"
    volumes:
      - n8n_data:/home/node/.n8n
    networks:
      - agent_net

volumes:
  ollama_data:
  qdrant_data:
  n8n_data:

networks:
  agent_net:
    driver: bridge

Agent Dockerfile

FROM python:3.11-slim

WORKDIR /app

# Dependencies
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# Agent code
COPY . .

# Health check
HEALTHCHECK --interval=30s --timeout=10s --retries=3 \
    CMD python -c "import requests; requests.get('http://localhost:8000/health')"

# Start
CMD ["python", "agent.py"]

Practical example: agent in Docker

# agent.py
import os
import requests
from fastapi import FastAPI

app = FastAPI()

OLLAMA_URL = os.getenv("OLLAMA_URL", "http://ollama:11434")
QDRANT_URL = os.getenv("QDRANT_URL", "http://qdrant:6333")

@app.get("/health")
def health():
    return {"status": "ok"}

@app.post("/agent")
async def run_agent(task: dict):
    """Agent endpoint"""
    # Agent logic here
    response = requests.post(f"{OLLAMA_URL}/api/chat", json={
        "model": "llama3.1",
        "messages": [{"role": "user", "content": task["task"]}],
        "tools": tools,
        "stream": False
    })
    return response.json()

Multi-container communication

# Agent communicates with Ollama over the Docker network
OLLAMA_URL = "http://ollama:11434"  # Service name as hostname

# Not localhost! Docker networking uses service names
# ollama = container name
# 11434 = port inside the container

Security considerations

  • Container isolation: Each container is isolated. A compromised container doesn’t affect others.
  • Networking: Keep internal communication on the Docker network, don’t expose it unnecessarily.
  • Secrets: Never include secrets in the image. Use environment variables or a secrets management system.
  • Resource limits: Set CPU and memory limits so one container can’t consume all resources.
  • Updates: Pull new images regularly and rebuild containers with the latest versions.

Common pitfalls

  • Localhost vs. service names: Containers communicate via service names, not localhost.
  • GPU in containers: For GPU support, use nvidia-docker or —gpus all. Not straightforward.
  • Data persistence: Container data is lost on restart. Use volumes for persistent data.
  • Networking: Containers need a shared network to communicate.
  • Logs: Container logs can grow large. Configure log rotation.

Further reading

Key Takeaways:

  • Docker = containers for each agent component.
  • Docker Compose orchestrates Ollama, agent, database, and vector store.
  • Isolated, reproducible, portable.
  • Use service names for container-to-container communication.
  • Docker is the standard for production deployments.

FAQ

Why use Docker for agents?

Isolation, reproducibility, and portability. Each component runs in its own container, all working as a unified stack. Deploy everything with docker-compose up.

What do I need for Docker agents?

Docker and Docker Compose. A Dockerfile for the agent and a docker-compose.yml for the stack (Ollama, agent, database, vector store).

How do I use GPU with Docker?

Use nvidia-docker or —gpus all. In docker-compose, add deploy.resources.reservations.devices with driver: nvidia.

How do containers communicate?

Via service names on the Docker network: http://ollama:11434 instead of http://localhost:11434. Docker resolves service names to container IP addresses.

Will my data persist after a restart?

Yes, with volumes. Container data is lost on restart, but volumes persist it. Use volumes for models, databases, and configurations.

Can I scale?

Yes, with docker-compose scale agent=3 or Kubernetes for large deployments. For horizontal scaling, run multiple containers behind a load balancer.

Docker or Kubernetes?

Docker for small to medium deployments. Kubernetes for large, complex deployments with auto-scaling, self-healing, and advanced orchestration.

Is Docker secure?

Yes, with best practices: container isolation, no secrets in images, resource limits, regular updates, and minimal base images.

Sources and further reading

Back to Blog
Share:

Related Posts