Ollama with Docker: Containerized Local AI
What this article covers
- How to run Ollama as a Docker container, with and without GPU support
- What volumes and port mapping mean for your container
- How to create a reproducible configuration with docker-compose
- Common pitfalls with GPU support and data persistence
- How to use Ollama in Docker on Windows, macOS, and Linux
Introduction: Ollama with Docker explained
Ollama is a runtime environment for local language models. You download a model and run it directly on your machine. If you’re already familiar with Ollama, you know that installing it on Linux, macOS, or Windows is straightforward. Sometimes, though, that’s not enough. Maybe you want to run Ollama on a server, a NAS device, or simply keep it isolated from the rest of your system. That’s where Docker comes in.
Docker packages Ollama in a container, an isolated environment with everything the tool needs to run. You start the container, use the model, and stop it again. No conflicts with other programs, no scattered files on your system. This article walks you through setting up Ollama with Docker step by step, from your first docker pull to a working docker-compose configuration.
If you’re new to Ollama, start with the Ollama overview and the Ollama installation guide. For Docker basics, see Docker Fundamentals.
Why use Docker for Ollama?
Imagine you’re already running multiple services on your server. A web server is running, along with a database and maybe a media server. Now you want to add Ollama. Without Docker, you’d install it directly on the system. Ollama creates files in system directories, sets up a service, and configures the network stack. When you want to remove it later, traces often remain behind.
Docker changes that picture. You pull the official Ollama image, start a container, and everything runs in isolation. When you need to update Ollama, just pull the latest image and restart the container. Your configuration stays intact because it lives in a volume. You can stop, delete, and recreate the container without losing your models.
Docker also gives you a reproducible configuration. You define how Ollama should run in a docker-compose.yml file, and you can replicate that same setup on any machine or server. This is especially valuable if you run Ollama on multiple machines or want to share your configuration with others.
Ollama with Docker at a glance
You pull the official ollama/ollama image from Docker Hub, start a container with port mapping for the API, and attach a volume to persist your models. For GPU support, you add the NVIDIA Container Toolkit and pass the GPU to the container using --gpus all. With docker-compose, you can describe the entire setup in a single file and start it with one command.
Who is Docker with Ollama for?
Docker with Ollama is for anyone who wants more control and cleanliness than a direct installation provides. That includes server administrators managing multiple services on one machine and wanting to avoid conflicts. It includes NAS owners running Ollama on their Synology or Ugreen devices. And it includes developers who need a reproducible environment for testing and integration.
If you just want to try Ollama briefly on your laptop, a direct installation is simpler. Once you need long-term operation, multiple models, or server setups, Docker is the better choice.
Key terms for Ollama and Docker
| Term | Definition |
|---|---|
| Docker | Platform that runs applications in isolated containers |
| Container | Running instance of an image, isolated from the host system |
| Image | Template for containers, containing all files and settings |
| Volume | Persistent storage that preserves data beyond the container’s lifecycle |
| docker-compose | Tool for defining and running multiple containers via a YAML file |
| GPU Runtime | Docker extension that gives containers access to the graphics card |
| NVIDIA Container Toolkit | Software package that makes Nvidia GPUs available in Docker containers |
| Port Mapping | Forwarding a host port to the container so services are accessible from outside |
| Persistence | The property that data survives even after the container stops or is deleted |
| Image Tag | Version identifier for an image, for example ollama/ollama:latest |
Prerequisites
Before you start, you need Docker on your system. On Linux, install Docker via the official installer or your distribution’s package manager. On Windows and macOS, use Docker Desktop. Details are in Docker Fundamentals.
If you want to run Ollama with GPU support, you’ll also need:
- A current Nvidia driver on the host system
- The NVIDIA Container Toolkit so Docker can access the GPU
On CPU-only setups, Docker alone is sufficient. For more GPU configuration details, see GPU Support in Docker. For an overview of CPU vs. GPU operation, visit CPU vs. GPU.
Step 1: Pull the Ollama image
Download the official Ollama image from Docker Hub:
docker pull ollama/ollama
By default, Docker pulls the latest tag. To specify a particular version:
docker pull ollama/ollama:0.3.0
The image contains the Ollama binary and all dependencies. It’s the foundation for every container you create afterward.
Step 2: Start your first container (CPU)
The simplest CPU-only startup looks like this:
docker run -d --name ollama -p 11434:11434 ollama/ollama
Here’s what each parameter does:
-d: Runs the container in the background (detached mode)--name ollama: Names the containerollamaso you can reference it easily-p 11434:11434: Maps port 11434 on the host to port 11434 in the containerollama/ollama: The image on which the container is based
After startup, verify the container is running:
docker ps
The Ollama API is now accessible at http://localhost:11434. Learn more in Using the Ollama API.
Step 3: Container with GPU support
For GPU support, you first need the NVIDIA Container Toolkit. On Ubuntu, install it like this:
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Then start Ollama with GPU support:
docker run -d --gpus all --name ollama-gpu -p 11434:11434 ollama/ollama
The --gpus all parameter passes all available GPUs to the container. If you have multiple GPUs and want to use only one, assign it explicitly:
docker run -d --gpus '"device=0"' --name ollama-gpu -p 11434:11434 ollama/ollama
Check whether the GPU is recognized in the container with:
docker exec -it ollama-gpu nvidia-smi
If nvidia-smi displays the GPU, everything is configured correctly. For more GPU configuration details, see GPU Support in Docker.
Step 4: Persistence with Volumes
Without a volume, your downloaded models disappear as soon as you delete the container. A volume stores models outside the container, so they persist across restarts and container recreation.
docker run -d --gpus all --name ollama-gpu -p 11434:11434 -v ollama-data:/root/.ollama ollama/ollama
The -v ollama-data:/root/.ollama flag creates a named volume called ollama-data and mounts it in the container at /root/.ollama. Ollama stores models, configurations, and metadata there.
You can also bind a local directory instead of a named volume:
docker run -d --gpus all --name ollama-gpu -p 11434:11434 -v /home/user/ollama:/root/.ollama ollama/ollama
This gives you direct access to files on the host system, which is convenient for backups and inspection.
Step 5: docker-compose.yml
For reproducible setups, docker-compose is the way to go. You describe the container in a YAML file and start everything with a single command.
Create a file named docker-compose.yml:
services:
ollama:
image: ollama/ollama:latest
container_name: ollama
ports:
- "11434:11434"
volumes:
- ollama-data:/root/.ollama
restart: unless-stopped
# Uncomment the following lines to enable GPU support:
# deploy:
# resources:
# reservations:
# devices:
# - driver: nvidia
# count: all
# capabilities: [gpu]
volumes:
ollama-data:
Start the container with:
docker compose up -d
Stop and remove it with:
docker compose down
The ollama-data volume persists. To delete it as well, use docker compose down -v.
To enable GPU support, uncomment the deploy entries. In older Docker versions, you can use runtime: nvidia instead, but this approach is now deprecated.
Step 6: Loading a Model into the Container
Once the container is running, load a model into it using docker exec to run commands inside the running container:
docker exec -it ollama ollama pull llama3
Then start the model:
docker exec -it ollama ollama run llama3
An interactive chat opens directly in your terminal. Type /bye to exit the session.
Since you’ve mounted a volume, the model persists when you restart the container. You won’t need to download it again.
Step 7: Accessing the API from Outside
The -p 11434:11434 flag in your docker run command, or the ports entry in docker-compose, makes the Ollama API accessible from your host system. By default, Ollama listens on 0.0.0.0:11434 inside the container.
Test the connection from your host:
curl http://localhost:11434/
To reach Ollama from other devices on your network, bind the port to your host’s network IP. In docker-compose:
ports:
- "0.0.0.0:11434:11434"
Pay attention to firewall rules and authentication. For more on secure network access, see Ollama Network Access.
Docker on Windows and macOS
On Windows and macOS, use Docker Desktop. It provides a graphical interface and WSL2 backend support on Windows. You can use the same docker run commands and docker-compose files as on Linux.
On Windows with WSL2 backend, GPU support works if you have an Nvidia GPU and WSL2 GPU drivers installed. The container command remains unchanged:
docker run -d --gpus all --name ollama -p 11434:11434 -v ollama-data:/root/.ollama ollama/ollama
On macOS with Apple Silicon, Docker Desktop automatically uses Metal acceleration. The --gpus all flag is not needed on macOS and can be ignored. Apple GPUs are provided through the host system, not through the NVIDIA Container Toolkit.
One limitation on Windows and macOS: Docker runs in a virtual machine, which can cause slight performance overhead. For maximum performance on a dedicated server, native Linux is the best choice.
Common Pitfalls with Ollama and Docker
GPU not recognized in the container. Often the NVIDIA Container Toolkit is missing or the Docker runtime wasn’t reconfigured. Check with nvidia-smi on the host and docker exec -it ollama nvidia-smi inside the container.
Models disappear after restart. You didn’t mount a volume. Without -v, Ollama stores models only in the container filesystem, which is lost when you delete the container.
Port 11434 already in use. If Ollama is already installed directly on your host, it’s using the same port. Stop the local Ollama service or map the container to a different port, such as -p 11435:11434.
Container fails to start after image update. If you pull a new image without removing the old container, the old version keeps running. Stop the container, remove it, and restart it. Your volume keeps your models safe.
OOM Killer terminates the container. Large models consume significant RAM. If your host doesn’t have enough memory, the Linux kernel kills the container. Check RAM usage with docker stats, choose a smaller model, or add more RAM.
docker-compose ignores GPU settings. In newer Docker versions, runtime: nvidia no longer works. Use the deploy.resources.reservations.devices block as shown in the example above.
Firewall blocks port 11434. If you want to reach Ollama from another machine, ensure your host’s firewall allows port 11434. On Ubuntu with UFW, use sudo ufw allow 11434.
Hardware, Costs, and Security with Ollama and Docker
Hardware requirements depend on the model you want to run. A 7B model like Llama 3 runs on 8 GB RAM, ideally with a GPU that has at least 6 GB VRAM. Larger models like Llama 3 70B need significantly more resources. Learn more in CPU vs. GPU and Running LLMs Locally.
Costs come from the hardware itself. Ollama and Docker are free and open source. No licensing fees or API costs apply as long as you run everything locally.
For security, don’t expose port 11434 to the internet unprotected. Use a reverse proxy with authentication if you want to make Ollama accessible remotely. Details are in Ollama Network Access. In Docker, you can also use network isolation by defining a custom Docker network and exposing only the necessary ports.
Further Reading and Resources on Ollama with Docker
- Ollama Overview for a general introduction
- Ollama Installation for direct installation without Docker
- Ollama API for the HTTP interface
- Ollama Configuration for environment variables and paths
- Ollama Network Access for secure remote access
- Running LLMs Locally for fundamentals on local models
- Docker Basics for getting started with Docker
- GPU Support in Docker for GPU configuration
- CPU vs. GPU for comparing CPU and GPU operation
FAQ: Ollama with Docker - Common Questions
Do I need Docker to run Ollama?
No. Ollama can be installed directly on Linux, macOS, and Windows. Docker is optional and works best for servers, NAS devices, and isolated environments.
Does Ollama work in Docker without a GPU?
Yes. Without the --gpus all parameter, Ollama runs on CPU. It’s slower, but sufficient for smaller models and testing.
Where does Docker store Ollama models?
Inside the container at /root/.ollama. If you mount a volume, models are stored in the volume and persist even after the container is removed.
How do I update Ollama in Docker?
Pull the latest image with docker pull ollama/ollama, stop the running container, remove it, and restart it. Your volume preserves models and configuration.
Can I run multiple Ollama containers simultaneously?
Yes. Give each container its own name and port, for example -p 11435:11434 for a second instance. Use separate volumes if the containers should hold different models.
How do I use Ollama with Docker on a Synology NAS?
Synology provides a graphical Docker interface in Container Manager. You can load the ollama/ollama image there and configure it with the same parameters as on the command line. Ensure you have sufficient RAM and CPU capacity.
What’s the difference between named volumes and bind mounts?
Named volumes are managed by Docker and stored in the Docker data directory. Bind mounts attach any directory from the host and give you direct file access. Both ensure persistence.
Does GPU support work on Windows with Docker Desktop?
Yes, with the WSL2 backend and appropriate Nvidia drivers. The --gpus all parameter works the same as on Linux. On macOS with Apple Silicon, the GPU is automatically used via Metal.
Can I run Open WebUI alongside Ollama in Docker?
Yes. Define both services in a docker-compose.yml and connect them through a shared Docker network. Open WebUI then reaches the Ollama container via its container name.
How do I ensure Ollama starts automatically after a reboot?
Use restart: unless-stopped in docker-compose or the --restart unless-stopped parameter with docker run. Docker will automatically restart the container when the Docker service starts.
Can I run Ollama in Docker with Podman?
Yes. Podman is largely compatible with Docker commands and docker-compose. The syntax for volumes and ports is identical. For GPU support, use --device nvidia.com/gpu=all with Podman.
References and Further Reading
- Ollama Docker Hub Repository
- Ollama Official Documentation
- Docker Documentation on Volumes and Port Mapping
- NVIDIA Container Toolkit Installation Guide
- Docker Compose Specification


