GPU Monitoring for Local AI
What this article covers
- Why GPU monitoring matters for local AI.
- Which metrics to track.
- Available tools for NVIDIA, AMD, and Intel.
- How to capture utilization, VRAM, temperature, and power consumption.
- Building dashboards, setting alerts, and avoiding common pitfalls.
Introduction: GPU monitoring for local AI
Local AI workloads place heavy demands on your graphics card. Large models fill VRAM, computations generate heat, and sluggish responses often signal resource constraints. Setting up GPU monitoring lets you spot bottlenecks early, prevent crashes, and make informed decisions about upgrades or optimization.
GPU monitoring isn’t just for large servers. Even a single mini-PC or workstation benefits from visibility into how much memory different models need and how temperature and performance evolve throughout the day.
Why do I need GPU monitoring?
Without it, you’re operating blind:
- You can’t tell whether your model runs on GPU or CPU.
- VRAM shortages cause crashes.
- High temperatures damage hardware over time.
- Slow workflows remain unexplained.
- Multi-user setups become hard to plan.
With monitoring, you can:
- Right-size models to your hardware,
- Choose appropriate quantization levels,
- Adjust fans and cooling,
- Catch usage spikes before they cause problems,
- Estimate costs and power draw.
Key metrics
- GPU utilization: The percentage of GPU actively computing right now.
- VRAM usage: How much video memory is in use, measured in GB.
- GPU temperature: Heat output of the card.
- Power consumption: Current draw in watts.
- Clock speed: GPU and memory frequencies.
- Per-process VRAM: Which process is consuming how much memory.
- Time-to-first-token: How quickly the model generates its first response.
- Tokens per second: Generation speed during inference.
Tools for NVIDIA
nvidia-smi
The standard command-line tool for NVIDIA GPUs. Reports utilization, VRAM, temperature, and running processes.
nvidia-smi
nvidia-smi dmon
nvidia-smi pmon
nvtop
A top-like interface for GPUs. More readable than nvidia-smi and also supports AMD and Intel.
nvtop
Prometheus and Grafana
For long-term monitoring and dashboards. NVIDIA Data Center GPU Manager or community exporters feed metrics into Prometheus.
Netdata
A lightweight agent-based monitoring solution with built-in GPU metrics.
Tools for AMD
- radeontop: Shows AMD GPU utilization.
- amdgpu-pro: Proprietary tools and drivers.
- ROCm-SMI: Similar to nvidia-smi for the ROCm platform.
- nvtop: Supports AMD.
Tools for Intel
- intel-gpu-top: Intel GPU utilization.
- xpu-smi: For Intel Data Center GPUs.
- nvtop: Supports Intel Arc and integrated GPUs.
Capturing metrics in Docker
When Ollama or vLLM run in containers, you need to read GPU metrics from the host or mount access inside the container. NVIDIA Container Toolkit makes nvidia-smi available inside containers. Prometheus exporters typically run as separate containers.
Building dashboards
A good dashboard displays:
- Current GPU utilization and VRAM per card.
- Historical load and temperature over 24 hours.
- Which processes are using the GPU.
- Alerts when temperature climbs too high or VRAM fills up.
- Comparison across multiple GPUs.
Alerts and thresholds
Reasonable alert levels:
- Temperature: Alert at 80 degrees Celsius.
- VRAM: Warning at 85 percent utilization, critical at 95 percent.
- GPU utilization: Sustained 95-100 percent is normal, but watch temperature.
- Processes: Unexpected processes consuming excessive VRAM.
Common pitfalls
- No monitoring inside containers: GPU data invisible from within the container.
- Only snapshots: You miss spikes and trends.
- Wrong tools for the hardware: AMD and Intel require different tools than NVIDIA.
- Alerts configured too late: Thresholds set too high or misconfigured entirely.
- Missing process attribution: You can’t tell which service is stressing the GPU.
Further reading and resources
FAQ: GPU monitoring
Which tool is the easiest starting point? For NVIDIA, nvidia-smi is already available. nvtop is cleaner and works across all vendors.
Can I display GPU metrics in Grafana? Yes, using Prometheus exporters for NVIDIA, AMD, or Intel.
What does high GPU utilization with low VRAM mean? Your model fits in VRAM, but the GPU is working hard on computation.
How do I spot VRAM bottlenecks? VRAM usage near capacity combined with slow or interrupted generation are telltale signs.
Should I prioritize temperature or utilization? Both matter. High utilization at reasonable temperatures is normal. Sustained high temperature degrades hardware over time.
Sources and further reading
- nvtop: https://github.com/Syllo/nvtop
- nvidia-smi docs: https://developer.nvidia.com/nvidia-management-library-nvml
- Netdata: https://www.netdata.cloud/
Summary: GPU monitoring for local AI
GPU monitoring keeps local AI workloads stable and efficient. Track GPU utilization, VRAM, temperature, and clock speed. Tools like nvidia-smi, nvtop, Prometheus, and Grafana cover most scenarios. For shared systems and always-on deployments, long-term monitoring and alerts are essential. Keeping your GPU in view helps you choose better models, avoid bottlenecks, and extend hardware life.


