GPU Passthrough in Docker
What this article covers
- How to pass GPUs to Docker containers.
- Differences between Nvidia and AMD setups.
- Configuration for Ollama, PyTorch, and other AI tools.
- ROCm, CUDA, and container runtimes.
- Common errors and how to fix them.
Introduction: GPU Passthrough in Docker
AI applications like Ollama, vLLM, and PyTorch benefit significantly from GPU acceleration. Docker lets you pass a host’s GPU directly to containers, allowing multiple GPU workloads to run isolated from one another. Nvidia and AMD have different requirements, but both work with Docker once you have the right drivers and runtimes installed.
This article shows how to make Nvidia and AMD GPUs available to Docker containers and how to avoid common pitfalls.
Key terms
- GPU Passthrough: Passing a physical GPU into a container.
- CUDA: Nvidia’s platform for parallel computing.
- ROCm: AMD’s open-source platform for GPU computing.
- Container Runtime: The execution environment that runs the container.
- Nvidia Container Toolkit: An extension that enables Nvidia GPUs in Docker.
- Device Mapping: Assigning device files to a container.
- Privileged: Container with full host access.
- Compute-Only: GPU used for computation only.
Nvidia GPU in Docker
Prerequisites
- Official Nvidia drivers installed.
- Docker and Docker Compose installed.
- Nvidia Container Toolkit.
Install Nvidia Container Toolkit
distribution=$(. /etc/os-release;echo $ID$VERSION_ID)
curl -s -L https://nvidia.github.io/libnvidia-container/gpgkey | sudo apt-key add -
curl -s -L https://nvidia.github.io/libnvidia-container/$distribution/libnvidia-container.list | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
sudo systemctl restart docker
Ollama with GPU
services:
ollama:
image: ollama/ollama
container_name: ollama
runtime: nvidia
environment:
- NVIDIA_VISIBLE_DEVICES=all
volumes:
- ollama-data:/root/.ollama
restart: unless-stopped
volumes:
ollama-data:
Or using the docker run command:
docker run -d --name ollama --gpus all -v ollama-data:/root/.ollama ollama/ollama
Verify GPU usage
docker exec -it ollama nvidia-smi
AMD GPU in Docker
Prerequisites
- AMD GPU with ROCm support.
- ROCm drivers installed.
- Compatible Linux distribution, usually Ubuntu.
Ollama with ROCm
services:
ollama:
image: ollama/ollama:rocm
container_name: ollama
devices:
- /dev/kfd
- /dev/dri
environment:
- HSA_OVERRIDE_GFX_VERSION=11.0.0
volumes:
- ollama-data:/root/.ollama
restart: unless-stopped
volumes:
ollama-data:
Verify GPU usage
docker exec -it ollama rocm-smi
Select a specific GPU
docker run --gpus '"device=0,1"' -d ollama/ollama
In Compose:
environment:
- NVIDIA_VISIBLE_DEVICES=0,1
Performance issues
- Wrong runtime: Container does not use
nvidia. - Driver mismatch: Nvidia driver and toolkit versions don’t match.
- ROCm version: Not every AMD GPU is officially supported.
- Memory limit: Container doesn’t have access to enough VRAM.
- CPU fallback: Ollama falls back to CPU if GPU is not detected.
- Mesa conflicts: Graphics drivers can collide with compute drivers.
Security
--gpus allexposes all GPUs.- Assign individual GPUs explicitly.
- Use privileged mode only when absolutely necessary.
- Restrict container images to trusted sources.
Further reading and resources
- BotServ.de Ollama Commands
- BotServ.de Ollama Performance
- BotServ.de Docker Security
- BotServ.de Ollama Performance
FAQ: GPU Passthrough
Do I need a GPU for Ollama? No, but it’s highly recommended for large models or responsive answers.
Does every Nvidia GPU work? Almost all current Nvidia GPUs with CUDA support do.
Can I use my AMD GPU?
Yes, if it supports ROCm or works with environment variables like HSA_OVERRIDE_GFX_VERSION.
Do I need to install the host driver inside the container? No, the host driver is passed through via the Container Toolkit.
How do I verify the GPU is detected in the container?
Run nvidia-smi or rocm-smi inside the container.
Sources and further reading
- Nvidia Container Toolkit: https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/
- ROCm Docker: https://rocmdocs.amd.com/
- Ollama Docker: https://hub.docker.com/r/ollama/ollama
Summary: GPU Passthrough in Docker
GPU passthrough in Docker lets you run Ollama and other AI tools on dedicated hardware without impacting the host operating system. Nvidia requires the Container Toolkit, while AMD needs ROCm and the correct device files. The key is using the right runtimes, matching driver versions, and assigning GPUs explicitly. If nvidia-smi or rocm-smi runs successfully inside your container, you know the GPU is properly exposed.


