Skip to content
BotServBotServ
MonitoringPrometheusGrafanaDockerOllamaSelf-HostingOperations

Monitoring for Local AI Systems

Monitor Ollama, Docker, hardware and AI services. Metrics, tools, alerts and best practices for reliable operation.

S

schutzgeist

5 min read
Monitoring for Local AI Systems

Monitoring for Local AI Systems

What This Article Covers

  • Why monitoring is essential for AI systems.
  • Which metrics matter for GPU, CPU, RAM, and services.
  • How Prometheus, Grafana, and Docker Stats work together.
  • Building alerts and dashboards for Ollama and AI agents.
  • Common pitfalls and best practices.

Introduction: Monitoring Local AI Systems

Running local AI systems means understanding what happens behind the scenes. A model that suddenly responds slowly, a container that crashes, or a GPU that overheats are problems you want to catch early. Monitoring collects data about hardware, services, and requests, then displays it in dashboards or triggers alerts.

AI systems behave differently than traditional web applications. They experience load spikes, long response times, and heavy memory demands. A server running at 20 percent capacity during text requests might hit its limits instantly when generating images. Without monitoring, you won’t spot these patterns.

Why Do You Need Monitoring?

Monitoring is more than a nice extra. It’s the foundation of stable operations. Without data on CPU, GPU, RAM, network, and services, you’re always reacting after problems occur. You only intervene once something visibly slows down or fails. With monitoring, you catch bottlenecks before they affect users.

Concrete benefits:

  • Early warnings: Alerts when load spikes, services go down, or storage fills up.
  • Cost control: Visualize power consumption, GPU usage, and resource allocation.
  • Error analysis: Logs and metrics reveal why a service crashed.
  • Optimization: Identify which models or workloads consume the most resources.
  • Planning: Make upgrade decisions based on data.

How Monitoring Works

Monitoring consists of several layers:

  • Metrics: Numerical values over time, like CPU usage as a percentage.
  • Logs: Text output from applications and systems.
  • Traces: Journey data of individual requests through multiple services.
  • Dashboards: Visual representation of metrics.
  • Alerts: Notifications when thresholds are crossed.

For home server or small office setups, metrics and logs usually suffice. Traces become relevant with distributed agent systems.

Who Should Use Monitoring?

  • Home server operators wanting stability.
  • Developers running AI services continuously.
  • Teams wanting to prevent outages and bottlenecks.
  • Users tracking GPU costs and power consumption.

Key Monitoring Concepts

  • Prometheus: Open source system for collecting and storing metrics.
  • Grafana: Visualization tool for dashboards.
  • Exporter: Small tools that format metrics from services for Prometheus.
  • Node Exporter: Collects hardware and system metrics on Linux.
  • cAdvisor: Provides metrics for Docker containers.
  • Time Series: Sequences of measurements over time.
  • Alertmanager: Sends alerts from Prometheus.

Critical Metrics for Local AI

Hardware Metrics

  • CPU usage: Percentage of processor utilization.
  • CPU temperature: Especially important under sustained load.
  • RAM usage: Total and available memory.
  • GPU usage: Computational load on the graphics card.
  • VRAM usage: Video memory in use.
  • GPU temperature: Critical during training and with large models.
  • Disk usage: Model and data sizes grow quickly.
  • Network traffic: Relevant especially with cloud integrations.

Service Metrics

  • Ollama running: Is the process active?
  • API response time: How long does a request take?
  • Container status: Are Docker containers running?
  • Error rate: How many requests fail?
  • Model usage: Which models are loaded and how long are they running?
  • Response length: How many tokens does the model output?

Application Metrics

  • Requests per hour: Shows usage spikes.
  • Queue wait time: Important in multi-user deployments.
  • Success rate: How often did the system deliver a useful result?
  • Memory per workspace: For AnythingLLM, Open WebUI, and RAG systems.

Practical Example: Setting Up Prometheus and Grafana

1. Install Prometheus

Prometheus starts quickly with Docker:

services:
  prometheus:
    image: prom/prometheus:latest
    container_name: prometheus
    ports:
      - "9090:9090"
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml
      - prometheus_data:/prometheus
    restart: unless-stopped

volumes:
  prometheus_data:

The prometheus.yml file defines scrape targets:

scrape_configs:
  - job_name: 'node'
    static_configs:
      - targets: ['node-exporter:9100']
  - job_name: 'cadvisor'
    static_configs:
      - targets: ['cadvisor:8080']

2. Node Exporter for Hardware Metrics

  node-exporter:
    image: prom/node-exporter:latest
    container_name: node-exporter
    ports:
      - "9100:9100"
    restart: unless-stopped

3. cAdvisor for Docker Containers

  cadvisor:
    image: gcr.io/cadvisor/cadvisor:latest
    container_name: cadvisor
    privileged: true
    devices:
      - /dev/kmsg:/dev/kmsg
    volumes:
      - /:/rootfs:ro
      - /var/run:/var/run:ro
      - /sys:/sys:ro
      - /var/lib/docker:/var/lib/docker:ro
    ports:
      - "8080:8080"
    restart: unless-stopped

4. Grafana for Dashboards

  grafana:
    image: grafana/grafana:latest
    container_name: grafana
    ports:
      - "3000:3000"
    volumes:
      - grafana_data:/var/lib/grafana
    restart: unless-stopped

volumes:
  grafana_data:

After startup, visit http://localhost:9090 for Prometheus and http://localhost:3000 for Grafana.

Monitoring Ollama

Ollama doesn’t expose Prometheus metrics natively, but you can monitor:

  • Process status: Is Ollama running?
  • API calls: curl http://localhost:11434/api/tags shows loaded models.
  • Resource consumption: cAdvisor displays CPU and RAM for the Ollama container.
  • GPU metrics: NVIDIA DCGM or nvidia-smi provide VRAM and GPU load.

For NVIDIA GPUs, consider the NVIDIA DCGM Exporter or script nvidia-smi output to a file that Prometheus reads.

Setting Up Meaningful Alerts

Alerts should only fire for real problems. Too many alarms lead to alert fatigue. Practical threshold examples:

  • RAM over 90 percent: Risk of OOM kill.
  • GPU temperature over 85 degrees: Risk of thermal throttling or damage.
  • Ollama unreachable: Service is down.
  • Disk over 85 percent full: Storage bottleneck imminent.
  • API response time over 30 seconds: Model or hardware overloaded.

Common Monitoring Pitfalls

  • Collecting too many metrics: Gather only what matters.
  • No retention policy: Prometheus defaults to 15 days. Configure longer periods if needed.
  • Missing labels: Without labels, you can’t distinguish between containers or services.
  • Scrape intervals too short: Frequent polling stresses the system.
  • Alerts without escalation: Define who gets notified and how.
  • Dashboards only: Visual monitoring doesn’t help with outages at 3 AM.

Further Reading and Resources

FAQ: Monitoring Local AI Systems

Do I need monitoring for a single small machine? Not strictly for testing. Once the system runs continuously or hosts multiple services, it becomes worthwhile.

What do Prometheus and Grafana cost? Both are open source and free. Costs come only from hardware, electricity, and storage.

Can I view logs with Grafana? Yes, with Loki. It centralizes log collection and search.

How often should Prometheus scrape data? 15 to 60 seconds is standard. Shorter intervals increase system load, longer intervals miss spikes.

Is Grafana necessary if I have Prometheus? No, but Grafana makes data much more readable and provides dashboards.

Can I run monitoring in Docker? Yes, Prometheus, Grafana, cAdvisor, and Node Exporter are all available as containers.

Sources and Further Reading

Summary: Monitoring Local AI Systems

Monitoring is essential for stable local AI systems. Prometheus collects metrics, Grafana visualizes them, and alerts notify you of problems. Focus on CPU, RAM, GPU, VRAM, container status, and API response times. A few containers create a solid monitoring infrastructure. Catching bottlenecks and outages early prevents surprises and enables targeted optimization.

Back to Blog
Share:

Related Posts