Skip to content
BotServBotServ
Error AnalysisDebuggingAI AgentsMonitoringTroubleshooting

Error Analysis for AI Agents

Error analysis for AI agents. Common errors, debugging, tracing, logging and practical examples for reliable agents.

S

schutzgeist

8 min read
Error Analysis for AI Agents

Error Analysis for AI Agents

What This Article Covers

  • Common errors in AI agents and how to spot them.
  • Setting up tracing, debugging, and logging for agents.
  • How to analyze and fix errors systematically.
  • Real-world examples of infinite loops, hallucinations, and tool failures.
  • Best practices for reliable agents.

Introduction: Error Analysis for AI Agents Explained

AI agents are intricate systems that invoke models, execute tools, and make decisions. When something breaks, an agent can get stuck in infinite loops, call the wrong tools, hallucinate facts, or crash entirely. Error analysis is the discipline of identifying, understanding, and fixing these problems.

This article is for developers building and debugging AI agents. You should already be familiar with what AI agents are and how to run them locally with Ollama. For Python fundamentals, see IRC-Coding.de.

Why Error Analysis Matters

Imagine your agent is supposed to sort emails, but it does nothing. Why? Did the model fail to respond? Did the tool break? Is the agent stuck in a loop? Without error analysis, you’re guessing blindly. With it, you see every step, every decision, every failure, and can intervene with precision.

Error Analysis for AI Agents in a Nutshell

Error analysis is the systematic investigation of agent failures. You use tracing (every step recorded), logging (events documented), and debugging (targeted investigation) to find the root cause and fix it.

The core principle: without tracing, there’s no debugging; without debugging, there’s no reliability.

Who Should Read This

  • Developers building and debugging AI agents.
  • System administrators managing agents in production.
  • Teams working to systematically improve agent quality.
  • Researchers investigating agent behavior.

You’ll need experience with Python, AI agents, and Ollama.

Key Terms

  • Tracing - Recording every agent step. Use it for: observability.
  • Logging - Structured logs. Use it for: foundation for analysis.
  • Debugging - Systematic error investigation. Use it for: finding root causes.
  • Infinite loop - Agent repeats steps without progress. Use it for: detecting a common error.
  • Hallucination - Model invents facts. Use it for: detecting a common error.
  • Tool error - Tool call fails. Use it for: detecting a common error.
  • AI agents - What you debug. Use it for: the subject of investigation.
  • Ollama - Local model server. Use it for: backend whose errors you analyze.
  • Function Calling - Tool use. Use it for: common error source.

Common Agent Errors

1. Infinite Loops

The agent repeats the same steps without making progress.

Symptoms: Agent runs for a very long time, identical tool calls repeat, no final answer.

Causes:

  • Model cannot solve the task and keeps retrying.
  • Tool returns unexpected results that the model misunderstands.
  • No step limit in place.

Fix: Add a step limit.

MAX_STEPS = 20
step = 0
while step < MAX_STEPS:
    step += 1
    # agent step
    if is_done(result):
        break
else:
    log_error("endless_loop", f"Agent exceeded {MAX_STEPS} steps")

2. Hallucinations

The model invents facts not found in the source material.

Symptoms: Answer contains unsourced claims, citations are incorrect.

Causes:

  • Model lacks sufficient context.
  • Model is prone to hallucinations (model-dependent).
  • No validation of outputs.

Fix: Validate facts and require sources.

See Research Workflows for strategies against hallucinations.

3. Tool Errors

Tool calls fail.

Symptoms: Tool returns an error, agent doesn’t understand it, execution stops.

Causes:

  • Tool is unreachable (network down, service offline).
  • Wrong parameters (model sent incorrect parameters).
  • Permission issue (tool execution not allowed).

Fix: Add error handling in the agent.

def call_tool_safe(tool_name, parameters):
    try:
        result = execute_tool(tool_name, parameters)
        return {"success": True, "result": result}
    except NetworkError as e:
        return {"success": False, "error": "tool_unreachable", "message": str(e)}
    except InvalidParameters as e:
        return {"success": False, "error": "invalid_parameters", "message": str(e)}
    except PermissionError as e:
        return {"success": False, "error": "permission_denied", "message": str(e)}

4. Context Length Exceeded

The conversation history grows too long and the model can no longer respond.

Symptoms: Model response is truncated, error about context length.

Causes:

  • Too many messages in history.
  • Long tool results.
  • No summarization mechanism.

Fix: Use summary-based memory management.

def manage_context(messages, max_tokens=4000):
    total = sum(len(m["content"]) for m in messages)
    if total > max_tokens:
        # Summarize early messages
        summary = call_ollama([
            {"role": "system", "content": "Summarize concisely."},
            {"role": "user", "content": str(messages[:5])}
        ])
        messages = [messages[0], {"role": "system", "content": f"Summary: {summary}"}] + messages[-3:]
    return messages

5. Model Unreachable

Ollama or your cloud API is not responding.

Symptoms: Timeout, connection error, agent stops.

Causes:

  • Ollama is not running.
  • Network connectivity issue.
  • Cloud API is down.

Fix: Implement retry logic and fallback.

import time

def call_model_retry(messages, max_retries=3, delay=1):
    for attempt in range(max_retries):
        try:
            return call_ollama(messages)
        except (ConnectionError, TimeoutError) as e:
            if attempt < max_retries - 1:
                time.sleep(delay * (attempt + 1))
            else:
                log_error("model_unreachable", str(e))
                raise

Implementing Tracing

from datetime import datetime
import json

class AgentTracer:
    def __init__(self, trace_file="agent_trace.json"):
        self.trace_file = trace_file
        self.steps = []

    def trace_step(self, step_type, data):
        step = {
            "timestamp": datetime.now().isoformat(),
            "step_type": step_type,
            "data": data
        }
        self.steps.append(step)
        with open(self.trace_file, "a") as f:
            f.write(json.dumps(step) + "\n")

    def trace_model_call(self, model, messages, response, duration_ms):
        self.trace_step("model_call", {
            "model": model,
            "input_messages": len(messages),
            "response_length": len(response),
            "duration_ms": duration_ms
        })

    def trace_tool_call(self, tool, parameters, result, duration_ms):
        self.trace_step("tool_call", {
            "tool": tool,
            "parameters": parameters,
            "result_status": result.get("status", "unknown"),
            "duration_ms": duration_ms
        })

    def trace_decision(self, decision, reasoning):
        self.trace_step("decision", {
            "decision": decision,
            "reasoning": reasoning
        })

    def trace_error(self, error_type, message, context):
        self.trace_step("error", {
            "error_type": error_type,
            "message": message,
            "context": context
        })

Agent with Tracing

class TraceableAgent:
    def __init__(self, model="llama3.1"):
        self.model = model
        self.tracer = AgentTracer()

    def run(self, task):
        self.tracer.trace_step("start", {"task": task})
        step = 0

        while step < 20:
            step += 1
            self.tracer.trace_step("step", {"step_number": step})

            # Call model
            start = time.time()
            try:
                response = call_ollama([
                    {"role": "user", "content": task}
                ])
                duration = (time.time() - start) * 1000
                self.tracer.trace_model_call(self.model, [task], response, duration)
            except Exception as e:
                self.tracer.trace_error("model_error", str(e), {"step": step})
                raise

            # Tool call?
            if has_tool_call(response):
                tool_name, params = extract_tool_call(response)
                start = time.time()
                try:
                    result = execute_tool(tool_name, params)
                    duration = (time.time() - start) * 1000
                    self.tracer.trace_tool_call(tool_name, params, result, duration)
                except Exception as e:
                    self.tracer.trace_error("tool_error", str(e), {"tool": tool_name})
                    result = {"success": False, "error": str(e)}

            # Done?
            if is_done(response):
                self.tracer.trace_step("done", {"result": response})
                return response

        self.tracer.trace_error("max_steps", "Maximum step count exceeded", {"steps": step})

Analyzing Traces

# All errors
jq 'select(.step_type == "error")' agent_trace.json

# All tool calls
jq 'select(.step_type == "tool_call")' agent_trace.json

# All failed tool calls
jq 'select(.step_type == "tool_call" and .data.result_status == "failed")' agent_trace.json

# Slow model calls
jq 'select(.step_type == "model_call" and .data.duration_ms > 5000)' agent_trace.json

Real-World Example 1: Debugging Infinite Loops

# Trace shows: Agent calls the same tool 20 times
# Analysis: Tool returns unexpected result, model doesn't understand it

# Solution: Provide clearer error descriptions to the model
def call_tool_with_explanation(tool_name, parameters):
    result = execute_tool(tool_name, parameters)
    if not result["success"]:
        # Clear error message for the model
        result["explanation"] = f"Tool {tool_name} failed: {result['error']}. Try a different approach."
    return result

Real-World Example 2: Debugging Hallucinations

# Trace shows: Model states facts without sources
# Analysis: Model lacks sufficient context

# Solution: Require sources
system_prompt = """
You are a research agent.
Every fact must have a source.
If you don't have a source, say "I don't know."
Do not invent facts.
"""

Real-World Example 3: Debugging Tool Errors

# Trace shows: Tool call fails with "connection refused"
# Analysis: Ollama is not running

# Solution: Health check before agent startup
def check_ollama_health():
    try:
        requests.get("http://localhost:11434/api/tags", timeout=5)
        return True
    except:
        return False

if not check_ollama_health():
    print("Ollama is not running. Starting Ollama.")
    subprocess.Popen(["ollama", "serve"])
    time.sleep(5)

Common Pitfalls

  • No tracing: Without traces, you can’t understand what went wrong.
  • No error handling: If a tool fails, the agent stops immediately.
  • No step limit: Agent can get stuck in infinite loops.
  • No context management: Context grows too long, model can’t respond.
  • No retry logic: A single network failure stops the agent.
  • Traces too verbose: Too many details make traces unreadable. Use log levels.

Further Reading

Key Takeaways:

  • Common issues: infinite loops, hallucinations, tool errors, context length.
  • Tracing records every step for full visibility.
  • Error handling catches tool failures without stopping the agent.
  • Step limits prevent infinite loops.
  • Retry logic and fallbacks improve reliability.

FAQ

What are the most common mistakes with AI agents?

Infinite loops (agent repeats steps), hallucinations (model invents facts), tool errors (calls fail), exceeded context length, and unreachable models.

What is tracing?

Tracing records every agent step: model calls, tool calls, decisions, errors. It lets you understand what happened.

How do I fix infinite loops?

Add a step limit (e.g., max 20 steps). When the agent reaches the limit, stop and log the error.

How do I fix hallucinations?

Require sources for every fact, validate facts against sources, use a system prompt that forbids hallucinations, and choose a model with lower hallucination rates.

How do I fix tool errors?

Implement error handling that distinguishes different error types (network, parameters, permissions). Give the model clear error descriptions so it can adapt.

How do I fix exceeded context length?

Use summary memory: when context gets too long, summarize old messages. Keep the system prompt and recent messages.

How do I make agents more reliable?

Implement retry logic for network errors, fall back to a local model if cloud services fail, and run health checks before starting the agent.

How do I analyze traces?

Use jq for JSON traces (filter by errors, tool calls, slow calls). For more complex analysis, use Python or tools like ELK Stack.

What do I do if Ollama is unreachable?

Use a health check to verify Ollama is running. If not, restart it. Implement retry logic to handle temporary outages.

How do I make agents production-ready?

Tracing, error handling, step limits, retry logic, context management, health checks, and audit logging. See Logging for details.

Sources and further reading

Back to Blog
Share:

Related Posts