Skip to content
BotServBotServ
OllamaDockerContainerGPUCompose

Run Ollama in Docker

Install Ollama as Docker container. GPU support, volumes, ports, and Compose configuration.

S

schutzgeist

3 min read
Run Ollama in Docker

Running Ollama in Docker

What this article covers

  • Why run Ollama in Docker.
  • Basic Docker setup.
  • GPU passthrough for NVIDIA and AMD.
  • Docker Compose configuration.
  • Tips for persistence and updates.

Introduction: Running Ollama in Docker

While Ollama is often installed directly on the host, it works equally well in a Docker container. This approach simplifies updates, isolation, and reproducibility. With GPU passthrough, Docker can run Ollama on your graphics card just like a native installation. For homelabs and servers, Docker provides a clean way to manage Ollama.

This article shows how to install and configure Ollama in Docker.

Key terms

  • Container: Isolated runtime environment.
  • Volume: Persistent storage.
  • GPU passthrough: Making the GPU available inside the container.
  • Compose: Declaratively define multiple containers.
  • NVIDIA Container Toolkit: Tools for GPU support in Docker.
  • runtime: Container runtime for GPU.
  • .ollama: Ollama data directory.
  • Port mapping: Forward a host port to the container.

Basic Docker command

docker run -d \
  -v ollama:/root/.ollama \
  -p 11434:11434 \
  --name ollama \
  ollama/ollama

With NVIDIA GPU

Requirement: NVIDIA Container Toolkit installed.

docker run -d \
  --gpus all \
  -v ollama:/root/.ollama \
  -p 11434:11434 \
  --name ollama \
  ollama/ollama

With AMD GPU

docker run -d \
  --device /dev/kfd \
  --device /dev/dri \
  -v ollama:/root/.ollama \
  -p 11434:11434 \
  --name ollama \
  ollama/ollama:rocm

Docker Compose

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    ports:
      - "11434:11434"
    volumes:
      - ollama-data:/root/.ollama
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

volumes:
  ollama-data:

Docker Compose with NVIDIA GPU

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    ports:
      - "11434:11434"
    volumes:
      - ollama-data:/root/.ollama
    runtime: nvidia
    environment:
      - NVIDIA_VISIBLE_DEVICES=all

volumes:
  ollama-data:

Docker Compose with AMD

services:
  ollama:
    image: ollama/ollama:rocm
    container_name: ollama
    ports:
      - "11434:11434"
    volumes:
      - ollama-data:/root/.ollama
    devices:
      - /dev/kfd
      - /dev/dri

volumes:
  ollama-data:

Download a model

docker exec -it ollama ollama pull llama3.1

Test the setup

curl http://localhost:11434/api/tags

Updates

docker pull ollama/ollama:latest
docker stop ollama
docker rm ollama
docker compose up -d

Or simply use Compose:

docker compose pull
docker compose up -d

Backing up data

The ollama-data volume contains your models and configuration:

docker run --rm -v ollama-data:/data -v $(pwd):/backup alpine \
  tar -czf /backup/ollama-backup.tar.gz -C /data .

Advantages

  • Easy updates.
  • Reproducible deployment.
  • Isolation from the host.
  • Simple integration into Compose stacks.
  • Good scalability on servers.

Limitations

  • GPU passthrough requires correct drivers and toolkit.
  • Device mounting is not always straightforward.
  • Pay attention to volume permissions.

Tips

  • Keep the volume persistent across restarts.
  • Don’t store models in the container layer.
  • Restart the container instead of recreating it when possible.
  • Use specific tags instead of latest in production.
  • Check logs regularly.
  • Install GPU drivers on the host beforehand.

Common pitfalls

  • GPU not detected: NVIDIA Container Toolkit or ROCm missing.
  • GPU not visible: Use deploy.resources instead of runtime: nvidia with newer Docker Compose versions.
  • Volume deleted: Models are lost.
  • Permission issues: Container runs as root.
  • Latest images: latest can introduce unexpected changes.

FAQ: Ollama in Docker

Do I need a GPU? No, but Ollama runs slower on CPU without one.

How do I start Ollama in Docker? Using ollama serve, which the image handles automatically.

Where are models stored? In the /root/.ollama volume.

Can I run multiple Ollama containers? Yes, on different ports with separate data volumes.

Which image for AMD? ollama/ollama:rocm.

Sources and further reading

Running Ollama in Docker is a clean, reproducible way to use language models locally or on a server. With GPU passthrough and a persistent volume, it works nearly as well as a native installation. Updates, backups, and integration into larger Compose stacks are simpler. Set up the NVIDIA Container Toolkit or ROCm correctly, and you get GPU acceleration inside the container.

Back to Blog
Share:

Related Posts