Logging for AI Agents
What this article covers
- How to set up logging for AI agents.
- Which events should be logged: tool calls, model responses, errors, and decisions.
- How to implement structured logging, audit trails, and compliance.
- Practical examples using Python, JSON logs, and log analysis.
- Best practices for security, performance, and retention.
Introduction: logging for AI agents explained
AI agents make decisions, call tools, and process data. When something goes wrong, you need to understand what happened. Logging is the recording of all agent actions so you can analyze failures, demonstrate compliance, and improve quality.
This article is for developers building AI agents and setting up logging. You should already understand what AI agents are and how to run them locally with Ollama. Python basics are available on IRC-Coding.de.
Why do you need logging for AI agents?
Imagine your agent deleted an email it shouldn’t have. Why did it do that? Which tools did it call? Which model made which decision? Without logging, you can’t trace it. With logging, you see every step, every decision, every tool call.
Logging becomes essential when agents perform critical actions, when compliance evidence is required (GDPR, ISO 27001), or when you want to systematically improve agent behavior.
Logging for AI agents in a nutshell
Logging is the recording of all agent actions in structured logs. Every tool call, model response, decision, and error gets recorded with a timestamp, context, and result. Logs can later be analyzed to find bugs, prove compliance, and improve quality.
The core principle: without logs, there’s no traceability; without traceability, there’s no trust.
Who this article is for
- Developers building AI agents and setting up logging.
- System administrators maintaining agents and analyzing failures.
- Compliance officers who need audit trails.
- Teams wanting to systematically improve agent quality.
You should be familiar with Python, AI agents, and Ollama.
Key terms in logging
- Logging - recording of events. Useful for: traceability.
- Audit log - record of all security-relevant actions. Useful for: compliance.
- Structured logging - logs in JSON format. Useful for: machine processing.
- Log level - priority of logs (DEBUG, INFO, WARNING, ERROR). Useful for: filtering.
- Tool call - agent invokes a tool. Useful for: critical to auditing.
- AI agents - programs that solve tasks autonomously. Useful for: what gets logged.
- Ollama - local model server. Useful for: backend whose calls are logged.
- Function calling - structured AI responses. Useful for: what gets logged.
- Compliance - adherence to regulations. Useful for: reason for audit logs.
What should be logged?
Model calls
Every call to the language model should be logged:
import logging
import json
from datetime import datetime
logger = logging.getLogger("agent")
logger.setLevel(logging.INFO)
def log_model_call(model, messages, response, duration_ms):
log_entry = {
"timestamp": datetime.now().isoformat(),
"type": "model_call",
"model": model,
"input_messages": len(messages),
"input_tokens": count_tokens(messages),
"output_tokens": count_tokens(response),
"duration_ms": duration_ms,
"status": "success"
}
logger.info(json.dumps(log_entry))
Tool calls
Every tool call is critical for auditing:
def log_tool_call(tool_name, parameters, result, duration_ms, status):
log_entry = {
"timestamp": datetime.now().isoformat(),
"type": "tool_call",
"tool": tool_name,
"parameters": parameters,
"result_summary": str(result)[:200],
"duration_ms": duration_ms,
"status": status
}
logger.info(json.dumps(log_entry))
Decisions
Every decision made by the agent:
def log_decision(decision_type, reasoning, action):
log_entry = {
"timestamp": datetime.now().isoformat(),
"type": "decision",
"decision_type": decision_type,
"reasoning": reasoning,
"action": action
}
logger.info(json.dumps(log_entry))
Errors
Every error with context:
def log_error(error_type, message, context):
log_entry = {
"timestamp": datetime.now().isoformat(),
"type": "error",
"error_type": error_type,
"message": message,
"context": context
}
logger.error(json.dumps(log_entry))
Structured logging with Python
Setup
import logging
import json
from datetime import datetime
class JSONFormatter(logging.Formatter):
def format(self, record):
log_entry = {
"timestamp": datetime.now().isoformat(),
"level": record.levelname,
"logger": record.name,
"message": record.getMessage()
}
if hasattr(record, "extra"):
log_entry.update(record.extra)
return json.dumps(log_entry)
# Configure logger
logger = logging.getLogger("agent")
logger.setLevel(logging.DEBUG)
# File handler
file_handler = logging.FileHandler("/var/log/agent.log")
file_handler.setFormatter(JSONFormatter())
logger.addHandler(file_handler)
# Console handler
console_handler = logging.StreamHandler()
console_handler.setFormatter(JSONFormatter())
logger.addHandler(console_handler)
Usage
# Log model call
logger.info("Model called", extra={
"type": "model_call",
"model": "llama3.1",
"duration_ms": 1234
})
# Log tool call
logger.info("Tool called", extra={
"type": "tool_call",
"tool": "send_email",
"parameters": {"to": "user@example.com"}
})
# Log error
logger.error("Tool failed", extra={
"type": "error",
"tool": "send_email",
"error": "SMTP timeout"
})
Implementing an audit trail
An audit trail is a specialized form of logging that records all security-relevant actions.
class AuditLogger:
def __init__(self, log_file="/var/log/agent_audit.log"):
self.logger = logging.getLogger("audit")
self.logger.setLevel(logging.INFO)
handler = logging.FileHandler(log_file)
handler.setFormatter(JSONFormatter())
self.logger.addHandler(handler)
def log_action(self, agent_id, action, target, parameters, result, user_id=None):
entry = {
"timestamp": datetime.now().isoformat(),
"agent_id": agent_id,
"user_id": user_id,
"action": action,
"target": target,
"parameters": parameters,
"result": result,
"status": "success" if result.get("success") else "failed"
}
self.logger.info(json.dumps(entry))
# Usage
audit = AuditLogger()
audit.log_action(
agent_id="email_agent_1",
action="send_email",
target="user@example.com",
parameters={"subject": "Re: Support-Anfrage"},
result={"success": True, "message_id": "abc123"},
user_id="user_42"
)
Practical Example: Complete Agent with Logging
class LoggedAgent:
def __init__(self, model="llama3.1"):
self.model = model
self.logger = logging.getLogger("agent")
self.audit = AuditLogger()
def call_model(self, messages):
start = time.time()
try:
response = requests.post(
"http://localhost:11434/api/chat",
json={"model": self.model, "messages": messages, "stream": False}
).json()
duration = (time.time() - start) * 1000
self.logger.info("Modell aufgerufen", extra={
"type": "model_call",
"model": self.model,
"duration_ms": duration,
"input_messages": len(messages),
"status": "success"
})
return response
except Exception as e:
duration = (time.time() - start) * 1000
self.logger.error("Modell-Aufruf fehlgeschlagen", extra={
"type": "error",
"model": self.model,
"duration_ms": duration,
"error": str(e)
})
raise
def call_tool(self, tool_name, parameters):
start = time.time()
try:
result = execute_tool(tool_name, parameters)
duration = (time.time() - start) * 1000
self.logger.info("Tool aufgerufen", extra={
"type": "tool_call",
"tool": tool_name,
"duration_ms": duration,
"status": "success"
})
self.audit.log_action(
agent_id=self.agent_id,
action=tool_name,
target=parameters.get("target", "unknown"),
parameters=parameters,
result={"success": True, "data": str(result)[:200]}
)
return result
except Exception as e:
duration = (time.time() - start) * 1000
self.logger.error("Tool fehlgeschlagen", extra={
"type": "error",
"tool": tool_name,
"duration_ms": duration,
"error": str(e)
})
self.audit.log_action(
agent_id=self.agent_id,
action=tool_name,
target=parameters.get("target", "unknown"),
parameters=parameters,
result={"success": False, "error": str(e)}
)
raise
Log Analysis
Searching Logs
# All errors
jq 'select(.level == "ERROR")' /var/log/agent.log
# All tool calls
jq 'select(.type == "tool_call")' /var/log/agent.log
# All failed tool calls
jq 'select(.type == "tool_call" and .status == "failed")' /var/log/agent.log
# Slow model calls (>5 seconds)
jq 'select(.type == "model_call" and .duration_ms > 5000)' /var/log/agent.log
Analyzing with Python
import json
def analyze_logs(log_file):
with open(log_file) as f:
logs = [json.loads(line) for line in f]
# Statistics
total_calls = len([l for l in logs if l.get("type") == "model_call"])
failed_tools = len([l for l in logs if l.get("type") == "tool_call" and l.get("status") == "failed"])
errors = len([l for l in logs if l.get("level") == "ERROR"])
# Average model duration
model_durations = [l["duration_ms"] for l in logs if l.get("type") == "model_call"]
avg_duration = sum(model_durations) / len(model_durations) if model_durations else 0
return {
"total_model_calls": total_calls,
"failed_tool_calls": failed_tools,
"total_errors": errors,
"avg_model_duration_ms": avg_duration
}
Log Rotation
Logs can grow quickly. Use log rotation to manage file sizes:
import logging.handlers
handler = logging.handlers.RotatingFileHandler(
"/var/log/agent.log",
maxBytes=100*1024*1024, # 100 MB
backupCount=5
)
handler.setFormatter(JSONFormatter())
logger.addHandler(handler)
Security Considerations
- Never log sensitive data: Avoid logging passwords, API keys, or confidential user information.
- Protect log files: Ensure logs are append-only and cannot be overwritten or deleted.
- Set retention policies: Define how long logs should be kept (e.g., 90 days for audit logs).
- Control access: Only authorized personnel should have read access to logs.
- Anonymize personal data: Remove or mask personally identifiable information in logs.
- See also: Audit Logging, Data Protection.
Common Pitfalls
- Too many logs: DEBUG-level logging in production creates massive files. Use appropriate log levels.
- Unstructured logs: Plain text logs are hard to parse and analyze. Switch to JSON format.
- Logging sensitive data: Passwords or API keys in logs create a security vulnerability.
- No log rotation: Logs grow indefinitely without rotation. Use RotatingFileHandler.
- Missing error logs: If you only log successful operations, you cannot troubleshoot failures.
- No timestamps: Without timestamps, you cannot correlate events across your system.
Further Reading on Logging
- AI Agent Fundamentals - What AI agents are.
- Running AI Agents with Ollama - Deploy agents locally.
- Audit Logging - Security and accountability.
- Data Protection - GDPR and AI.
- Human Approval - Human-in-the-loop patterns.
- Configuring Guardrails - Safety guardrails for agents.
- Email Automation - Practical example with logging.
Key Takeaways:
- Logging records all agent actions for visibility and debugging.
- Log model calls, tool invocations, decisions, and errors.
- Use structured logging (JSON) for machine-readable analysis.
- Maintain audit trails for compliance and security.
- Exclude sensitive data from logs and implement log rotation.
FAQ: Logging for AI Agents - Common Questions
What should I log?
Why use structured logging?
What is an audit trail?
How should I handle sensitive data in logs?
What is log rotation?
How long should I keep logs?
Which log levels should I use?
How do I analyze logs?
How much performance overhead does logging add?
Do I need logging for compliance?
Sources and Further Reading
- Python Logging - Official Python logging documentation.
- jq - JSON processor for logs.
- ELK Stack - Log analysis platform.
- Grafana Loki - Log aggregation.


