Skip to content
BotServBotServ
DockerGPUNVIDIAContainer ToolkitCUDAOllamaSelf-Hosting

Docker GPU Support: NVIDIA Container Toolkit

Use GPU in Docker containers: install, configure NVIDIA Container Toolkit and run Ollama with GPU. Step-by-step guide.

S

schutzgeist

10 min read
Docker GPU Support: NVIDIA Container Toolkit

Docker GPU Support: NVIDIA Container Toolkit

What this article covers

  • How to install the NVIDIA Container Toolkit and configure Docker for GPU access
  • How to pass the GPU to a container using --gpus all
  • How to enable GPU support in docker-compose via deploy.resources
  • How to run Ollama with GPU acceleration in Docker
  • Common pitfalls with drivers, runtime, and compose files

Introduction: Understanding Docker GPU support

Docker isolates applications in containers. By default, a container has no access to your host system’s graphics card. For pure CPU workloads, that’s fine. But as soon as you want to run AI models, video encoding, or CUDA-based calculations in a container, you need the GPU. That’s where the NVIDIA Container Toolkit comes in.

The toolkit extends Docker to pass your host system’s NVIDIA driver into containers. You install it once, configure the Docker runtime, and then you can grant any container GPU access by adding the --gpus all parameter. Inside the container, nvidia-smi becomes available, and CUDA applications run with full GPU acceleration.

This article walks you through installation on Ubuntu, testing GPU access in a container, and setting up Ollama with GPU support. If you’re new to Docker, start with Docker Basics. For an overview of self-hosting topics, see Self-Hosting.

Why do I need GPU in Docker?

Imagine running Ollama in a Docker container on your server. Without GPU support, the model runs on CPU. A 7B model like Llama 3 takes about 30 to 60 seconds on a modern CPU to generate the first tokens, then just a few tokens per second afterward. That’s barely usable for chat.

With GPU, things change dramatically. An NVIDIA GPU with 8 GB VRAM generates 30 to 60 tokens per second with the same model, and the first tokens appear in under a second. The difference is massive. For more context, see CPU vs. GPU and GPU Offloading.

For Ollama to use the GPU inside Docker, you need to pass the graphics card into the container. This is called GPU passthrough. The NVIDIA Container Toolkit makes this possible by bridging the host driver and the container. Without it, the container won’t see any GPU, no matter what parameters you pass.

Docker GPU in a nutshell

Install the NVIDIA Container Toolkit, configure the Docker runtime with nvidia-ctk runtime configure --runtime=docker, and restart Docker. After that, pass --gpus all to any container that needs GPU access. In docker-compose, use the deploy.resources.reservations.devices block with the nvidia driver. Inside the container, verify the GPU is detected with nvidia-smi.

Who is this article for?

This article is for anyone running GPU-accelerated workloads in Docker containers. That includes self-hosters running Ollama or other AI tools on a server with an NVIDIA GPU. It covers developers testing and deploying CUDA applications in containers. And it helps administrators isolate and manage multiple GPU services on a single machine.

If you’ve never used Docker before, start with Docker Basics. If you’re new to Ollama, read Ollama Overview first.

Key terms

TermMeaning
NVIDIA Container ToolkitSoftware package that makes NVIDIA GPUs available in Docker containers
nvidia-smiCommand-line tool that displays GPU status, VRAM, and processes
CUDANVIDIA platform for parallel computing on the GPU
Docker RuntimeComponent that starts and manages containers, extensible via GPU support
—gpusDocker parameter that passes one or more GPUs to a container
GPU PassthroughPassing the physical GPU to a container or VM
ContainerIsolated runtime environment that executes an image with all dependencies
VRAMVideo memory of the GPU where models and intermediate results reside
DriverNVIDIA driver on the host system, required for GPU access
RuntimeExecution environment, configurable in Docker via nvidia-ctk

Prerequisites

Before you start, you need two things on your host system:

1. A working NVIDIA driver. Check that the GPU is recognized by running nvidia-smi on the host:

nvidia-smi

If you see a table with the GPU name, driver version, and VRAM, the driver is installed correctly. If the command is missing or returns an error, install the NVIDIA driver first. On Ubuntu, use:

sudo ubuntu-drivers autoinstall
sudo reboot

2. Docker installed. Check your Docker version:

docker --version

If Docker is missing, install it from the official installer. Instructions are available in Docker Basics and the Ubuntu article.

Installing NVIDIA Container Toolkit

On Ubuntu, installation takes just a few steps. You add the NVIDIA repository, install the package, and configure the Docker runtime.

Step 1: Add repository key

curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg

Step 2: Register repository

curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
  sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
  sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

Step 3: Install package

sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit

Step 4: Configure Docker runtime

sudo nvidia-ctk runtime configure --runtime=docker

This command adds the NVIDIA runtime to Docker’s configuration. You can verify it:

cat /etc/docker/daemon.json

The file should contain an entry for nvidia as the default runtime or as a named runtime.

Step 5: Restart Docker

sudo systemctl restart docker

Docker now picks up the new configuration. From now on, you can pass the GPU to containers.

Testing Docker with GPU

After installation, verify that the GPU is accessible in a container. Use the official NVIDIA CUDA image:

docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi

If the output shows the GPU table, everything is configured correctly. The container has access to the graphics card.

You can also pass a specific GPU if you have multiple:

docker run --rm --gpus '"device=0"' nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi

The parameter '"device=0"' selects the first GPU. Use '"device=0,1"' to pass two GPUs.

Ollama with GPU in Docker

Ollama benefits significantly from GPU acceleration. Once you have the toolkit installed, you can start Ollama with GPU support:

docker run -d --gpus all --name ollama-gpu -p 11434:11434 -v ollama-data:/root/.ollama ollama/ollama

Here’s what each parameter does:

  • -d: Run the container in the background
  • --gpus all: Pass all GPUs to the container
  • --name ollama-gpu: Container name for easy reference
  • -p 11434:11434: Port mapping for the Ollama API
  • -v ollama-data:/root/.ollama: Volume for model persistence

Verify that the GPU is recognized inside the container:

docker exec -it ollama-gpu nvidia-smi

If the GPU table appears, Ollama is using your graphics card. For more details on Ollama with Docker, see Ollama with Docker.

docker-compose with GPU

For reproducible setups, use docker-compose. GPU configuration happens via the deploy.resources block.

Create a docker-compose.yml file:

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama-gpu
    ports:
      - "11434:11434"
    volumes:
      - ollama-data:/root/.ollama
    restart: unless-stopped
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

volumes:
  ollama-data:

Start the container:

docker compose up -d

The deploy.resources.reservations.devices block tells Docker that the container needs an NVIDIA GPU. The count: all entry passes all GPUs. Use count: 1 to pass exactly one. The capabilities: [gpu] line marks the device type.

In older Docker versions, runtime: nvidia was used. This no longer works reliably in current versions. Use the deploy block instead.

Example: Ollama GPU in Docker

Here’s a complete example from installation to running a model.

Step 1: Install the toolkit (as described above)

Step 2: Create docker-compose.yml

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama-gpu
    ports:
      - "11434:11434"
    volumes:
      - ollama-data:/root/.ollama
    restart: unless-stopped
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

volumes:
  ollama-data:

Step 3: Start the container

docker compose up -d

Step 4: Verify GPU in the container

docker exec -it ollama-gpu nvidia-smi

Step 5: Download a model

docker exec -it ollama-gpu ollama pull llama3

Step 6: Run the model

docker exec -it ollama-gpu ollama run llama3

A chat interface opens in your terminal. GPU acceleration delivers fast responses. Type /bye to end the session.

Common Pitfalls

GPU not recognized in the container. Usually this means the runtime configuration is missing or Docker wasn’t restarted. Check with nvidia-smi on the host and docker exec -it ollama-gpu nvidia-smi in the container. If it works on the host but not in the container, you’re missing nvidia-ctk runtime configure --runtime=docker or the Docker restart.

NVIDIA driver is outdated. The toolkit requires a compatible driver. If your driver is too old, installation fails or nvidia-smi doesn’t work in the container. Update the driver and restart.

--gpus not recognized. This parameter exists only in Docker 19.03 and later. If you’re on an older version, upgrade Docker. Check with docker --version.

docker-compose ignores GPU settings. In current Docker versions, runtime: nvidia no longer works. Use the deploy.resources.reservations.devices block. Make sure you’re running docker-compose v2, not the old Python version.

Container runs but CUDA application crashes. Sometimes CUDA libraries are missing from the container. Use an image that already includes CUDA, such as nvidia/cuda:12.4.0-base-ubuntu22.04, or install CUDA in the Dockerfile.

VRAM runs out. Large models demand substantial VRAM. If the GPU doesn’t have enough video memory, the application crashes or falls back to CPU mode. Check VRAM usage with nvidia-smi while the model runs. Learn more at GPU-Offloading.

Multiple containers share one GPU. When several GPU containers run at the same time, they share VRAM. This can cause out-of-memory errors. Monitor VRAM with nvidia-smi and distribute workloads across multiple GPUs or start containers sequentially.

Hardware, Cost, and Security

Hardware: You need an NVIDIA GPU with sufficient VRAM. For 7B models, 8 GB VRAM is enough; for 13B models, 12 to 16 GB; for 70B models, considerably more. The host needs enough RAM and CPU to run the container and operating system. See CPU vs. GPU for hardware decisions.

Cost: The NVIDIA Container Toolkit is free and open source. Docker is also free. Costs come only from hardware: GPU, server, and electricity. No licensing fees apply.

Security: GPU passthrough gives the container full access to the GPU. This poses no security risk to the host as long as you use trustworthy images. Make sure to use only official images from Docker Hub or trusted sources. If you have multiple users on one server, GPU containers should not be exposed to the network without protection. Use firewall rules and, if possible, a VPN like Tailscale for remote access.

Further Reading

FAQ: Docker GPU Support

Do I need the NVIDIA Container Toolkit for AMD GPUs?

No. The toolkit is specific to NVIDIA. For AMD GPUs, use the ROCm stack and pass the GPU via --device /dev/kfd and --device /dev/dri. Configuration differs significantly from NVIDIA.

Does GPU support work without the NVIDIA driver on the host?

No. The toolkit forwards the host driver into the container. Without a working NVIDIA driver on the host, there is no GPU in the container. Install the driver first and verify it with nvidia-smi on the host.

Can I split multiple GPUs across different containers?

Yes. Use --gpus '"device=0"' to pass only the first GPU, or '"device=1"' for the second. In docker-compose, use count: 1 and add device_ids: [0] in the device block.

What does capabilities: [gpu] mean in docker-compose?

This entry marks the device type the container needs. gpu is the standard for GPU access. You can also specify compute, utility, graphics, or video if you need specific CUDA functions.

Does GPU support work on Windows with Docker Desktop?

Yes, with WSL2 backend and matching NVIDIA drivers for WSL. The --gpus all parameter works as on Linux. On macOS with Apple Silicon, the GPU is provided via Metal and the toolkit is not needed.

How do I see how much VRAM my container uses?

Run nvidia-smi on the host while the container is running. The table shows all processes using GPU resources and their VRAM consumption. Alternatively, run docker exec -it ollama-gpu nvidia-smi in the container.

Do I need to restart Docker after updating the NVIDIA driver?

Yes. After a driver update, restart the host or at least the Docker service with sudo systemctl restart docker. Otherwise, the container may access an outdated driver version.

Can I enable GPU support later without recreating the container?

No. The --gpus parameter is set when the container starts. You must stop the container, remove it, and start a new one with --gpus all. The volume keeps your data safe.

What’s the difference between --gpus all and runtime: nvidia?

--gpus all is the current standard since Docker 19.03. runtime: nvidia was the old way and is now deprecated. In docker-compose, use the deploy.resources block instead of the runtime line.

Do I need CUDA in the container if I’ve installed the toolkit?

It depends on your application. The toolkit forwards the driver, but CUDA libraries must be present in the image. Ollama includes them in its image. For custom applications, use a CUDA image or install CUDA in the Dockerfile.

Can I run Ollama with GPU in Docker on a NAS?

Yes, if the NAS has an NVIDIA GPU and runs Linux, like Ugreen or Asrock devices. Synology NAS typically use Intel or AMD and don’t directly support the NVIDIA Container Toolkit. Check your NAS hardware support.

References

Back to Blog
Share:

Related Posts