Skip to content
BotServBotServ
LoggingAudit LoggingAI AgentsMonitoringCompliance

Logging for AI Agents

Set up logging for AI agents. Audit logs, structured logging, tool calls, error analysis and compliance.

S

schutzgeist

8 min read
Logging for AI Agents

Logging for AI Agents

What this article covers

  • How to set up logging for AI agents.
  • Which events should be logged: tool calls, model responses, errors, and decisions.
  • How to implement structured logging, audit trails, and compliance.
  • Practical examples using Python, JSON logs, and log analysis.
  • Best practices for security, performance, and retention.

Introduction: logging for AI agents explained

AI agents make decisions, call tools, and process data. When something goes wrong, you need to understand what happened. Logging is the recording of all agent actions so you can analyze failures, demonstrate compliance, and improve quality.

This article is for developers building AI agents and setting up logging. You should already understand what AI agents are and how to run them locally with Ollama. Python basics are available on IRC-Coding.de.

Why do you need logging for AI agents?

Imagine your agent deleted an email it shouldn’t have. Why did it do that? Which tools did it call? Which model made which decision? Without logging, you can’t trace it. With logging, you see every step, every decision, every tool call.

Logging becomes essential when agents perform critical actions, when compliance evidence is required (GDPR, ISO 27001), or when you want to systematically improve agent behavior.

Logging for AI agents in a nutshell

Logging is the recording of all agent actions in structured logs. Every tool call, model response, decision, and error gets recorded with a timestamp, context, and result. Logs can later be analyzed to find bugs, prove compliance, and improve quality.

The core principle: without logs, there’s no traceability; without traceability, there’s no trust.

Who this article is for

  • Developers building AI agents and setting up logging.
  • System administrators maintaining agents and analyzing failures.
  • Compliance officers who need audit trails.
  • Teams wanting to systematically improve agent quality.

You should be familiar with Python, AI agents, and Ollama.

Key terms in logging

  • Logging - recording of events. Useful for: traceability.
  • Audit log - record of all security-relevant actions. Useful for: compliance.
  • Structured logging - logs in JSON format. Useful for: machine processing.
  • Log level - priority of logs (DEBUG, INFO, WARNING, ERROR). Useful for: filtering.
  • Tool call - agent invokes a tool. Useful for: critical to auditing.
  • AI agents - programs that solve tasks autonomously. Useful for: what gets logged.
  • Ollama - local model server. Useful for: backend whose calls are logged.
  • Function calling - structured AI responses. Useful for: what gets logged.
  • Compliance - adherence to regulations. Useful for: reason for audit logs.

What should be logged?

Model calls

Every call to the language model should be logged:

import logging
import json
from datetime import datetime

logger = logging.getLogger("agent")
logger.setLevel(logging.INFO)

def log_model_call(model, messages, response, duration_ms):
    log_entry = {
        "timestamp": datetime.now().isoformat(),
        "type": "model_call",
        "model": model,
        "input_messages": len(messages),
        "input_tokens": count_tokens(messages),
        "output_tokens": count_tokens(response),
        "duration_ms": duration_ms,
        "status": "success"
    }
    logger.info(json.dumps(log_entry))

Tool calls

Every tool call is critical for auditing:

def log_tool_call(tool_name, parameters, result, duration_ms, status):
    log_entry = {
        "timestamp": datetime.now().isoformat(),
        "type": "tool_call",
        "tool": tool_name,
        "parameters": parameters,
        "result_summary": str(result)[:200],
        "duration_ms": duration_ms,
        "status": status
    }
    logger.info(json.dumps(log_entry))

Decisions

Every decision made by the agent:

def log_decision(decision_type, reasoning, action):
    log_entry = {
        "timestamp": datetime.now().isoformat(),
        "type": "decision",
        "decision_type": decision_type,
        "reasoning": reasoning,
        "action": action
    }
    logger.info(json.dumps(log_entry))

Errors

Every error with context:

def log_error(error_type, message, context):
    log_entry = {
        "timestamp": datetime.now().isoformat(),
        "type": "error",
        "error_type": error_type,
        "message": message,
        "context": context
    }
    logger.error(json.dumps(log_entry))

Structured logging with Python

Setup

import logging
import json
from datetime import datetime

class JSONFormatter(logging.Formatter):
    def format(self, record):
        log_entry = {
            "timestamp": datetime.now().isoformat(),
            "level": record.levelname,
            "logger": record.name,
            "message": record.getMessage()
        }
        if hasattr(record, "extra"):
            log_entry.update(record.extra)
        return json.dumps(log_entry)

# Configure logger
logger = logging.getLogger("agent")
logger.setLevel(logging.DEBUG)

# File handler
file_handler = logging.FileHandler("/var/log/agent.log")
file_handler.setFormatter(JSONFormatter())
logger.addHandler(file_handler)

# Console handler
console_handler = logging.StreamHandler()
console_handler.setFormatter(JSONFormatter())
logger.addHandler(console_handler)

Usage

# Log model call
logger.info("Model called", extra={
    "type": "model_call",
    "model": "llama3.1",
    "duration_ms": 1234
})

# Log tool call
logger.info("Tool called", extra={
    "type": "tool_call",
    "tool": "send_email",
    "parameters": {"to": "user@example.com"}
})

# Log error
logger.error("Tool failed", extra={
    "type": "error",
    "tool": "send_email",
    "error": "SMTP timeout"
})

Implementing an audit trail

An audit trail is a specialized form of logging that records all security-relevant actions.

class AuditLogger:
    def __init__(self, log_file="/var/log/agent_audit.log"):
        self.logger = logging.getLogger("audit")
        self.logger.setLevel(logging.INFO)
        handler = logging.FileHandler(log_file)
        handler.setFormatter(JSONFormatter())
        self.logger.addHandler(handler)

    def log_action(self, agent_id, action, target, parameters, result, user_id=None):
        entry = {
            "timestamp": datetime.now().isoformat(),
            "agent_id": agent_id,
            "user_id": user_id,
            "action": action,
            "target": target,
            "parameters": parameters,
            "result": result,
            "status": "success" if result.get("success") else "failed"
        }
        self.logger.info(json.dumps(entry))

# Usage
audit = AuditLogger()
audit.log_action(
    agent_id="email_agent_1",
    action="send_email",
    target="user@example.com",
    parameters={"subject": "Re: Support-Anfrage"},
    result={"success": True, "message_id": "abc123"},
    user_id="user_42"
)

Practical Example: Complete Agent with Logging

class LoggedAgent:
    def __init__(self, model="llama3.1"):
        self.model = model
        self.logger = logging.getLogger("agent")
        self.audit = AuditLogger()

    def call_model(self, messages):
        start = time.time()
        try:
            response = requests.post(
                "http://localhost:11434/api/chat",
                json={"model": self.model, "messages": messages, "stream": False}
            ).json()
            duration = (time.time() - start) * 1000

            self.logger.info("Modell aufgerufen", extra={
                "type": "model_call",
                "model": self.model,
                "duration_ms": duration,
                "input_messages": len(messages),
                "status": "success"
            })
            return response
        except Exception as e:
            duration = (time.time() - start) * 1000
            self.logger.error("Modell-Aufruf fehlgeschlagen", extra={
                "type": "error",
                "model": self.model,
                "duration_ms": duration,
                "error": str(e)
            })
            raise

    def call_tool(self, tool_name, parameters):
        start = time.time()
        try:
            result = execute_tool(tool_name, parameters)
            duration = (time.time() - start) * 1000

            self.logger.info("Tool aufgerufen", extra={
                "type": "tool_call",
                "tool": tool_name,
                "duration_ms": duration,
                "status": "success"
            })

            self.audit.log_action(
                agent_id=self.agent_id,
                action=tool_name,
                target=parameters.get("target", "unknown"),
                parameters=parameters,
                result={"success": True, "data": str(result)[:200]}
            )
            return result
        except Exception as e:
            duration = (time.time() - start) * 1000
            self.logger.error("Tool fehlgeschlagen", extra={
                "type": "error",
                "tool": tool_name,
                "duration_ms": duration,
                "error": str(e)
            })
            self.audit.log_action(
                agent_id=self.agent_id,
                action=tool_name,
                target=parameters.get("target", "unknown"),
                parameters=parameters,
                result={"success": False, "error": str(e)}
            )
            raise

Log Analysis

Searching Logs

# All errors
jq 'select(.level == "ERROR")' /var/log/agent.log

# All tool calls
jq 'select(.type == "tool_call")' /var/log/agent.log

# All failed tool calls
jq 'select(.type == "tool_call" and .status == "failed")' /var/log/agent.log

# Slow model calls (>5 seconds)
jq 'select(.type == "model_call" and .duration_ms > 5000)' /var/log/agent.log

Analyzing with Python

import json

def analyze_logs(log_file):
    with open(log_file) as f:
        logs = [json.loads(line) for line in f]

    # Statistics
    total_calls = len([l for l in logs if l.get("type") == "model_call"])
    failed_tools = len([l for l in logs if l.get("type") == "tool_call" and l.get("status") == "failed"])
    errors = len([l for l in logs if l.get("level") == "ERROR"])

    # Average model duration
    model_durations = [l["duration_ms"] for l in logs if l.get("type") == "model_call"]
    avg_duration = sum(model_durations) / len(model_durations) if model_durations else 0

    return {
        "total_model_calls": total_calls,
        "failed_tool_calls": failed_tools,
        "total_errors": errors,
        "avg_model_duration_ms": avg_duration
    }

Log Rotation

Logs can grow quickly. Use log rotation to manage file sizes:

import logging.handlers

handler = logging.handlers.RotatingFileHandler(
    "/var/log/agent.log",
    maxBytes=100*1024*1024,  # 100 MB
    backupCount=5
)
handler.setFormatter(JSONFormatter())
logger.addHandler(handler)

Security Considerations

  • Never log sensitive data: Avoid logging passwords, API keys, or confidential user information.
  • Protect log files: Ensure logs are append-only and cannot be overwritten or deleted.
  • Set retention policies: Define how long logs should be kept (e.g., 90 days for audit logs).
  • Control access: Only authorized personnel should have read access to logs.
  • Anonymize personal data: Remove or mask personally identifiable information in logs.
  • See also: Audit Logging, Data Protection.

Common Pitfalls

  • Too many logs: DEBUG-level logging in production creates massive files. Use appropriate log levels.
  • Unstructured logs: Plain text logs are hard to parse and analyze. Switch to JSON format.
  • Logging sensitive data: Passwords or API keys in logs create a security vulnerability.
  • No log rotation: Logs grow indefinitely without rotation. Use RotatingFileHandler.
  • Missing error logs: If you only log successful operations, you cannot troubleshoot failures.
  • No timestamps: Without timestamps, you cannot correlate events across your system.

Further Reading on Logging

Key Takeaways:

  • Logging records all agent actions for visibility and debugging.
  • Log model calls, tool invocations, decisions, and errors.
  • Use structured logging (JSON) for machine-readable analysis.
  • Maintain audit trails for compliance and security.
  • Exclude sensitive data from logs and implement log rotation.

FAQ: Logging for AI Agents - Common Questions

What should I log?

Model calls (model name, duration, token count), tool invocations (tool name, parameters, results), decisions (type, reasoning, action), and errors (type, message, context).

Why use structured logging?

Structured logs in JSON format can be machine-parsed, filtered, and correlated easily. Unstructured text logs are difficult to search and analyze at scale.

What is an audit trail?

An audit trail is a complete record of all security-relevant actions. It shows who did what and when, providing accountability and enabling compliance verification.

How should I handle sensitive data in logs?

Never log passwords, API keys, or confidential user data. Anonymize personally identifiable information, and control access to log files strictly.

What is log rotation?

Log rotation limits log file size by archiving and compressing files once they reach a threshold, then deleting old backups after a specified count or period.

How long should I keep logs?

Retention depends on compliance requirements. Audit logs typically are kept for 90 days, while GDPR-sensitive logs may require longer retention. Define a clear retention policy.

Which log levels should I use?

Use DEBUG during development, INFO in production, WARNING for potential issues, and ERROR for failures. Production systems typically run at INFO or WARNING level.

How do I analyze logs?

Use jq for JSON log queries on the command line, Python for complex analysis, or dedicated tools like the ELK Stack or Grafana Loki for centralized log management.

How much performance overhead does logging add?

Structured logging has minimal performance impact when done asynchronously. Use async handlers to avoid blocking your agent while writing logs.

Do I need logging for compliance?

Yes. Compliance standards like GDPR and ISO 27001 require demonstrating who did what and when. Audit trails are essential for meeting these requirements.

Sources and Further Reading

Back to Blog
Share:

Related Posts