Docker GPU Support: NVIDIA Container Toolkit
What this article covers
- How to install the NVIDIA Container Toolkit and configure Docker for GPU access
- How to pass the GPU to a container using
--gpus all - How to enable GPU support in docker-compose via
deploy.resources - How to run Ollama with GPU acceleration in Docker
- Common pitfalls with drivers, runtime, and compose files
Introduction: Understanding Docker GPU support
Docker isolates applications in containers. By default, a container has no access to your host system’s graphics card. For pure CPU workloads, that’s fine. But as soon as you want to run AI models, video encoding, or CUDA-based calculations in a container, you need the GPU. That’s where the NVIDIA Container Toolkit comes in.
The toolkit extends Docker to pass your host system’s NVIDIA driver into containers. You install it once, configure the Docker runtime, and then you can grant any container GPU access by adding the --gpus all parameter. Inside the container, nvidia-smi becomes available, and CUDA applications run with full GPU acceleration.
This article walks you through installation on Ubuntu, testing GPU access in a container, and setting up Ollama with GPU support. If you’re new to Docker, start with Docker Basics. For an overview of self-hosting topics, see Self-Hosting.
Why do I need GPU in Docker?
Imagine running Ollama in a Docker container on your server. Without GPU support, the model runs on CPU. A 7B model like Llama 3 takes about 30 to 60 seconds on a modern CPU to generate the first tokens, then just a few tokens per second afterward. That’s barely usable for chat.
With GPU, things change dramatically. An NVIDIA GPU with 8 GB VRAM generates 30 to 60 tokens per second with the same model, and the first tokens appear in under a second. The difference is massive. For more context, see CPU vs. GPU and GPU Offloading.
For Ollama to use the GPU inside Docker, you need to pass the graphics card into the container. This is called GPU passthrough. The NVIDIA Container Toolkit makes this possible by bridging the host driver and the container. Without it, the container won’t see any GPU, no matter what parameters you pass.
Docker GPU in a nutshell
Install the NVIDIA Container Toolkit, configure the Docker runtime with nvidia-ctk runtime configure --runtime=docker, and restart Docker. After that, pass --gpus all to any container that needs GPU access. In docker-compose, use the deploy.resources.reservations.devices block with the nvidia driver. Inside the container, verify the GPU is detected with nvidia-smi.
Who is this article for?
This article is for anyone running GPU-accelerated workloads in Docker containers. That includes self-hosters running Ollama or other AI tools on a server with an NVIDIA GPU. It covers developers testing and deploying CUDA applications in containers. And it helps administrators isolate and manage multiple GPU services on a single machine.
If you’ve never used Docker before, start with Docker Basics. If you’re new to Ollama, read Ollama Overview first.
Key terms
| Term | Meaning |
|---|---|
| NVIDIA Container Toolkit | Software package that makes NVIDIA GPUs available in Docker containers |
| nvidia-smi | Command-line tool that displays GPU status, VRAM, and processes |
| CUDA | NVIDIA platform for parallel computing on the GPU |
| Docker Runtime | Component that starts and manages containers, extensible via GPU support |
| —gpus | Docker parameter that passes one or more GPUs to a container |
| GPU Passthrough | Passing the physical GPU to a container or VM |
| Container | Isolated runtime environment that executes an image with all dependencies |
| VRAM | Video memory of the GPU where models and intermediate results reside |
| Driver | NVIDIA driver on the host system, required for GPU access |
| Runtime | Execution environment, configurable in Docker via nvidia-ctk |
Prerequisites
Before you start, you need two things on your host system:
1. A working NVIDIA driver. Check that the GPU is recognized by running nvidia-smi on the host:
nvidia-smi
If you see a table with the GPU name, driver version, and VRAM, the driver is installed correctly. If the command is missing or returns an error, install the NVIDIA driver first. On Ubuntu, use:
sudo ubuntu-drivers autoinstall
sudo reboot
2. Docker installed. Check your Docker version:
docker --version
If Docker is missing, install it from the official installer. Instructions are available in Docker Basics and the Ubuntu article.
Installing NVIDIA Container Toolkit
On Ubuntu, installation takes just a few steps. You add the NVIDIA repository, install the package, and configure the Docker runtime.
Step 1: Add repository key
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
Step 2: Register repository
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
Step 3: Install package
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
Step 4: Configure Docker runtime
sudo nvidia-ctk runtime configure --runtime=docker
This command adds the NVIDIA runtime to Docker’s configuration. You can verify it:
cat /etc/docker/daemon.json
The file should contain an entry for nvidia as the default runtime or as a named runtime.
Step 5: Restart Docker
sudo systemctl restart docker
Docker now picks up the new configuration. From now on, you can pass the GPU to containers.
Testing Docker with GPU
After installation, verify that the GPU is accessible in a container. Use the official NVIDIA CUDA image:
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
If the output shows the GPU table, everything is configured correctly. The container has access to the graphics card.
You can also pass a specific GPU if you have multiple:
docker run --rm --gpus '"device=0"' nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
The parameter '"device=0"' selects the first GPU. Use '"device=0,1"' to pass two GPUs.
Ollama with GPU in Docker
Ollama benefits significantly from GPU acceleration. Once you have the toolkit installed, you can start Ollama with GPU support:
docker run -d --gpus all --name ollama-gpu -p 11434:11434 -v ollama-data:/root/.ollama ollama/ollama
Here’s what each parameter does:
-d: Run the container in the background--gpus all: Pass all GPUs to the container--name ollama-gpu: Container name for easy reference-p 11434:11434: Port mapping for the Ollama API-v ollama-data:/root/.ollama: Volume for model persistence
Verify that the GPU is recognized inside the container:
docker exec -it ollama-gpu nvidia-smi
If the GPU table appears, Ollama is using your graphics card. For more details on Ollama with Docker, see Ollama with Docker.
docker-compose with GPU
For reproducible setups, use docker-compose. GPU configuration happens via the deploy.resources block.
Create a docker-compose.yml file:
services:
ollama:
image: ollama/ollama:latest
container_name: ollama-gpu
ports:
- "11434:11434"
volumes:
- ollama-data:/root/.ollama
restart: unless-stopped
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
volumes:
ollama-data:
Start the container:
docker compose up -d
The deploy.resources.reservations.devices block tells Docker that the container needs an NVIDIA GPU. The count: all entry passes all GPUs. Use count: 1 to pass exactly one. The capabilities: [gpu] line marks the device type.
In older Docker versions, runtime: nvidia was used. This no longer works reliably in current versions. Use the deploy block instead.
Example: Ollama GPU in Docker
Here’s a complete example from installation to running a model.
Step 1: Install the toolkit (as described above)
Step 2: Create docker-compose.yml
services:
ollama:
image: ollama/ollama:latest
container_name: ollama-gpu
ports:
- "11434:11434"
volumes:
- ollama-data:/root/.ollama
restart: unless-stopped
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
volumes:
ollama-data:
Step 3: Start the container
docker compose up -d
Step 4: Verify GPU in the container
docker exec -it ollama-gpu nvidia-smi
Step 5: Download a model
docker exec -it ollama-gpu ollama pull llama3
Step 6: Run the model
docker exec -it ollama-gpu ollama run llama3
A chat interface opens in your terminal. GPU acceleration delivers fast responses. Type /bye to end the session.
Common Pitfalls
GPU not recognized in the container. Usually this means the runtime configuration is missing or Docker wasn’t restarted. Check with nvidia-smi on the host and docker exec -it ollama-gpu nvidia-smi in the container. If it works on the host but not in the container, you’re missing nvidia-ctk runtime configure --runtime=docker or the Docker restart.
NVIDIA driver is outdated. The toolkit requires a compatible driver. If your driver is too old, installation fails or nvidia-smi doesn’t work in the container. Update the driver and restart.
--gpus not recognized. This parameter exists only in Docker 19.03 and later. If you’re on an older version, upgrade Docker. Check with docker --version.
docker-compose ignores GPU settings. In current Docker versions, runtime: nvidia no longer works. Use the deploy.resources.reservations.devices block. Make sure you’re running docker-compose v2, not the old Python version.
Container runs but CUDA application crashes. Sometimes CUDA libraries are missing from the container. Use an image that already includes CUDA, such as nvidia/cuda:12.4.0-base-ubuntu22.04, or install CUDA in the Dockerfile.
VRAM runs out. Large models demand substantial VRAM. If the GPU doesn’t have enough video memory, the application crashes or falls back to CPU mode. Check VRAM usage with nvidia-smi while the model runs. Learn more at GPU-Offloading.
Multiple containers share one GPU. When several GPU containers run at the same time, they share VRAM. This can cause out-of-memory errors. Monitor VRAM with nvidia-smi and distribute workloads across multiple GPUs or start containers sequentially.
Hardware, Cost, and Security
Hardware: You need an NVIDIA GPU with sufficient VRAM. For 7B models, 8 GB VRAM is enough; for 13B models, 12 to 16 GB; for 70B models, considerably more. The host needs enough RAM and CPU to run the container and operating system. See CPU vs. GPU for hardware decisions.
Cost: The NVIDIA Container Toolkit is free and open source. Docker is also free. Costs come only from hardware: GPU, server, and electricity. No licensing fees apply.
Security: GPU passthrough gives the container full access to the GPU. This poses no security risk to the host as long as you use trustworthy images. Make sure to use only official images from Docker Hub or trusted sources. If you have multiple users on one server, GPU containers should not be exposed to the network without protection. Use firewall rules and, if possible, a VPN like Tailscale for remote access.
Further Reading
- Self-Hosting for an overview of the topic
- Docker Basics to get started with Docker
- Ollama with Docker for Ollama-specific Docker configuration
- Ollama Overview for a general introduction
- CPU vs. GPU to compare CPU and GPU operation
- GPU-Offloading for details on GPU usage
- Ubuntu for operating system basics
FAQ: Docker GPU Support
Do I need the NVIDIA Container Toolkit for AMD GPUs?
No. The toolkit is specific to NVIDIA. For AMD GPUs, use the ROCm stack and pass the GPU via --device /dev/kfd and --device /dev/dri. Configuration differs significantly from NVIDIA.
Does GPU support work without the NVIDIA driver on the host?
No. The toolkit forwards the host driver into the container. Without a working NVIDIA driver on the host, there is no GPU in the container. Install the driver first and verify it with nvidia-smi on the host.
Can I split multiple GPUs across different containers?
Yes. Use --gpus '"device=0"' to pass only the first GPU, or '"device=1"' for the second. In docker-compose, use count: 1 and add device_ids: [0] in the device block.
What does capabilities: [gpu] mean in docker-compose?
This entry marks the device type the container needs. gpu is the standard for GPU access. You can also specify compute, utility, graphics, or video if you need specific CUDA functions.
Does GPU support work on Windows with Docker Desktop?
Yes, with WSL2 backend and matching NVIDIA drivers for WSL. The --gpus all parameter works as on Linux. On macOS with Apple Silicon, the GPU is provided via Metal and the toolkit is not needed.
How do I see how much VRAM my container uses?
Run nvidia-smi on the host while the container is running. The table shows all processes using GPU resources and their VRAM consumption. Alternatively, run docker exec -it ollama-gpu nvidia-smi in the container.
Do I need to restart Docker after updating the NVIDIA driver?
Yes. After a driver update, restart the host or at least the Docker service with sudo systemctl restart docker. Otherwise, the container may access an outdated driver version.
Can I enable GPU support later without recreating the container?
No. The --gpus parameter is set when the container starts. You must stop the container, remove it, and start a new one with --gpus all. The volume keeps your data safe.
What’s the difference between --gpus all and runtime: nvidia?
--gpus all is the current standard since Docker 19.03. runtime: nvidia was the old way and is now deprecated. In docker-compose, use the deploy.resources block instead of the runtime line.
Do I need CUDA in the container if I’ve installed the toolkit?
It depends on your application. The toolkit forwards the driver, but CUDA libraries must be present in the image. Ollama includes them in its image. For custom applications, use a CUDA image or install CUDA in the Dockerfile.
Can I run Ollama with GPU in Docker on a NAS?
Yes, if the NAS has an NVIDIA GPU and runs Linux, like Ugreen or Asrock devices. Synology NAS typically use Intel or AMD and don’t directly support the NVIDIA Container Toolkit. Check your NAS hardware support.
References
- NVIDIA Container Toolkit Documentation: https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/
- Docker GPU Documentation: https://docs.docker.com/config/containers/resource_constraints/#gpu
- NVIDIA CUDA Docker Images: https://hub.docker.com/r/nvidia/cuda
- Ollama Docker Image: https://hub.docker.com/r/ollama/ollama


