Skip to content
BotServBotServ
DockerMonitoringGrafanaPrometheuscAdvisorNode Exporter

Docker Monitoring with Grafana and Prometheus

Monitor containers, resources, and AI services in Docker. Grafana, Prometheus, cAdvisor, and Node Exporter.

S

schutzgeist

4 min read
Docker Monitoring with Grafana and Prometheus

Docker Monitoring with Grafana and Prometheus

What this article covers

  • Key metrics for Docker containers.
  • Overview of Prometheus, Grafana, cAdvisor, and Node Exporter.
  • A Docker Compose example for a complete monitoring stack.
  • Dashboards, alerts, and best practices.
  • Common gotchas.

Introduction: Docker monitoring with Grafana and Prometheus

If you run Ollama, Open WebUI, and other services in Docker, you want to know how much CPU, RAM, network bandwidth, and storage they consume. A monitoring stack built from Prometheus, Grafana, cAdvisor, and Node Exporter displays these metrics clearly. This lets you spot bottlenecks early, before services fail or slow down.

This article walks you through setting up these components in Docker and which metrics matter most for AI workloads.

Key terms

  • Prometheus: Time-series database for metrics.
  • Grafana: Visualization tool for metrics and dashboards.
  • cAdvisor: Container Advisor, collects Docker container metrics.
  • Node Exporter: Collects host metrics such as CPU, RAM, and disk.
  • Scrape: Periodic collection of metrics from a target.
  • Target: An endpoint from which Prometheus fetches data.
  • Dashboard: Visual overview in Grafana.
  • Alert: Notification triggered when a threshold is breached.

Why monitoring matters

  • Spot bottlenecks before they become critical.
  • Track CPU and RAM usage per container.
  • Monitor GPU utilization for AI models.
  • Watch network and disk I/O.
  • Identify long-term trends and patterns.
  • Get alerted to problems as they happen.

Components

Prometheus

Prometheus stores metrics as time-series data. It queries defined endpoints at regular intervals and saves the values.

Grafana

Grafana connects to Prometheus and displays the data in dashboards. Pre-built dashboards exist for Docker, hosts, and GPUs.

cAdvisor

cAdvisor runs as a container and provides detailed metrics for all running Docker containers, including CPU, RAM, network, and filesystem statistics.

Node Exporter

Node Exporter runs directly on the host (or as a privileged container) and exposes hardware and operating system metrics.

Docker Compose example

services:
  prometheus:
    image: prom/prometheus:latest
    container_name: prometheus
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml
      - prometheus-data:/prometheus
    ports:
      - "127.0.0.1:9090:9090"
    command:
      - "--config.file=/etc/prometheus/prometheus.yml"
      - "--storage.tsdb.path=/prometheus"
    restart: unless-stopped

  grafana:
    image: grafana/grafana:latest
    container_name: grafana
    ports:
      - "127.0.0.1:3000:3000"
    volumes:
      - grafana-data:/var/lib/grafana
    restart: unless-stopped

  cadvisor:
    image: gcr.io/cadvisor/cadvisor:latest
    container_name: cadvisor
    privileged: true
    volumes:
      - /:/rootfs:ro
      - /var/run:/var/run:ro
      - /sys:/sys:ro
      - /var/lib/docker:/var/lib/docker:ro
      - /dev/disk:/dev/disk:ro
    ports:
      - "127.0.0.1:8080:8080"
    restart: unless-stopped

  node-exporter:
    image: prom/node-exporter:latest
    container_name: node-exporter
    volumes:
      - /proc:/host/proc:ro
      - /sys:/host/sys:ro
      - /:/rootfs:ro
    command:
      - "--path.procfs=/host/proc"
      - "--path.rootfs=/rootfs"
      - "--path.sysfs=/host/sys"
    restart: unless-stopped

volumes:
  prometheus-data:
  grafana-data:

Prometheus configuration

prometheus.yml:

global:
  scrape_interval: 15s

scrape_configs:
  - job_name: 'prometheus'
    static_configs:
      - targets: ['localhost:9090']

  - job_name: 'cadvisor'
    static_configs:
      - targets: ['cadvisor:8080']

  - job_name: 'node-exporter'
    static_configs:
      - targets: ['node-exporter:9100']

Grafana dashboards

Once running, log in at http://localhost:3000. Add Prometheus as a data source. Import pre-built dashboards using their IDs:

  • Docker: ID 11600
  • Node Exporter Full: ID 1860
  • cAdvisor: ID 14282

Key metrics for AI workloads

  • container_cpu_usage_seconds_total: CPU consumption per container.
  • container_memory_usage_bytes: RAM usage.
  • container_network_receive_bytes_total: Incoming network traffic.
  • container_fs_usage_bytes: Storage usage.
  • nvidia_gpu_utilization: GPU utilization (requires Nvidia Exporter).
  • node_memory_MemAvailable_bytes: Available RAM on the host.

Alerts

Grafana can send notifications when values cross thresholds. Common examples:

  • Container memory usage exceeds 90 percent of allocated RAM.
  • CPU load stays above 95 percent for an extended period.
  • Disk usage exceeds 85 percent.

GPU monitoring

For Nvidia GPUs, you can integrate the Nvidia DCGM Exporter or expose nvidia-smi metrics through Node Exporter. AMD GPUs require different tools such as rocm-smi or amdgpu-pro. See the GPU monitoring article for more detail.

Tips

  • Set up monitoring from the start, not as an afterthought.
  • Back up Prometheus and Grafana volumes regularly.
  • Keep monitoring behind a VPN or authentication layer if possible.
  • Set resource limits so monitoring doesn’t starve your system.
  • Watch long-term trends to plan for growth.

Common gotchas

  • cAdvisor not privileged: Container metrics won’t appear.
  • Wrong paths in Node Exporter: Metrics will be incomplete.
  • Prometheus can’t reach targets: Check container names and ports.
  • No data in Grafana: Verify the data source configuration.
  • Dashboard looks wrong: Metric names can vary between setups.
  • Running out of disk space: Prometheus data grows over time; configure retention.

Further reading

FAQ: Docker monitoring

Do I need all four components? For basic monitoring, Prometheus, Grafana, and cAdvisor suffice. Node Exporter adds host-level metrics.

How many resources does this monitoring stack use? Prometheus and Grafana are lightweight. cAdvisor can consume some CPU.

Can I monitor Nvidia GPUs? Yes, with the appropriate exporter. Check the GPU monitoring article.

Should I expose Grafana to the public internet? No. Keep it internal or behind a VPN.

How do I keep historical data long-term? Back up your volumes or configure Prometheus remote storage.

References

Summary: Docker monitoring with Grafana and Prometheus

A monitoring stack built from Prometheus, Grafana, cAdvisor, and Node Exporter gives you visibility into how hard your Docker containers and host are working. For AI workloads, CPU, RAM, network, storage, and GPU metrics are especially valuable. Pre-built dashboards and alerts help you catch bottlenecks early. The key is planning monitoring from day one, backing up your volumes, and securing access. That way, you maintain control over your homelab and AI infrastructure.

Back to Blog
Share:

Related Posts