Skip to content
BotServBotServ
Cost ControlTokenBudgetMonitoringLocal AI

Cost Control for AI Agents

Cost control for AI agents: token usage, API costs, local AI, budgets, alerts, and practical examples.

S

schutzgeist

8 min read
Cost Control for AI Agents

Cost Control for AI Agents

What this article covers

  • How to monitor token consumption and API costs for AI agents.
  • How to set up budgets, alerts, and limits.
  • How local AI eliminates API costs.
  • Practical examples for cost tracking, budget warnings, and optimization.
  • Best practices for cost-efficient agent operation.

Introduction: Understanding cost control for AI agents

AI agents can be expensive. Every agent step calls the language model, every call consumes tokens, and every token costs money when using cloud APIs. An agent running 24/7 that consumes 5 million tokens per day can cost thousands of euros per month. Cost control is the practice of monitoring, limiting, and optimizing this consumption.

This article is for developers who build AI agents and need to manage costs. You should understand what AI agents are and how to run them locally with Ollama. For Python fundamentals, see IRC-Coding.de.

Why do you need cost control?

Imagine you build an agent that sorts emails. It runs 24/7 and calls the model for each email. With 100 emails per day and 10,000 tokens per call, that’s 1 million tokens daily. Using GPT-4o (5 $/M input, 15 $/M output), you’re looking at $20 per day, $600 per month. Without cost control, you won’t notice until the bill arrives.

Cost control for AI agents explained

Cost control means monitoring and limiting the token consumption of AI agents. You measure how many tokens each agent consumes, set budgets, configure alerts, and optimize to reduce spending. With local AI (Ollama), API costs disappear entirely.

The core principle: what you don’t measure, you can’t control.

Who this article is for

  • Developers building AI agents who need to manage costs.
  • Teams running agents in production and needing to stick to budgets.
  • Decision makers who want to understand API costs.
  • Self-hosters evaluating local AI as a cost alternative.

Prior experience with AI agents and Ollama is helpful.

Key terms

  • Token - A unit of text, roughly 4 characters. Useful to know: it’s the billing unit for APIs.
  • AI agents - Programs that call models. Useful to know: what generates costs.
  • Ollama - Local model server. Useful to know: eliminates API costs.
  • Budget - Cost limit per time period. Useful to know: caps your spending.
  • Alert - Warning when budget is exceeded. Useful to know: lets you react early.
  • Quantization - Reducing model size. Useful to know: enables faster, cheaper inference.
  • Context length - Maximum tokens per call. Useful to know: affects cost.
  • Logging - Recording of calls. Useful to know: the foundation for cost tracking.

Measuring token consumption

Cloud API: tokens from the response

Cloud APIs return token counts in their response:

import requests

def call_model_with_tracking(messages):
    response = requests.post(
        "https://api.openai.com/v1/chat/completions",
        headers={"Authorization": "Bearer sk-..."},
        json={
            "model": "gpt-4o",
            "messages": messages,
            "stream": False
        }
    )
    data = response.json()
    return {
        "content": data["choices"][0]["message"]["content"],
        "input_tokens": data["usage"]["prompt_tokens"],
        "output_tokens": data["usage"]["completion_tokens"],
        "cost": calculate_cost("gpt-4o", data["usage"])
    }

def calculate_cost(model, usage):
    prices = {
        "gpt-4o": {"input": 5.0, "output": 15.0},  # per 1M tokens
        "gpt-4o-mini": {"input": 0.15, "output": 0.60},
        "claude-3-5-sonnet": {"input": 3.0, "output": 15.0}
    }
    p = prices.get(model, {"input": 0, "output": 0})
    return (usage["prompt_tokens"] * p["input"] + usage["completion_tokens"] * p["output"]) / 1_000_000

Ollama: tokens from the response

Ollama also returns token counts:

def call_ollama_with_tracking(messages):
    response = requests.post(
        "http://localhost:11434/api/chat",
        json={"model": "llama3.1", "messages": messages, "stream": False}
    )
    data = response.json()
    return {
        "content": data["message"]["content"],
        "input_tokens": data.get("prompt_eval_count", 0),
        "output_tokens": data.get("eval_count", 0),
        "cost": 0.0  # Local AI: no API costs
    }

Implementing a cost tracker

from datetime import datetime, timedelta
from collections import defaultdict
import json

class CostTracker:
    def __init__(self, log_file="agent_costs.json"):
        self.log_file = log_file
        self.costs = defaultdict(lambda: {"input_tokens": 0, "output_tokens": 0, "cost": 0.0, "calls": 0})

    def log_call(self, agent_id, model, input_tokens, output_tokens, cost):
        entry = {
            "timestamp": datetime.now().isoformat(),
            "agent_id": agent_id,
            "model": model,
            "input_tokens": input_tokens,
            "output_tokens": output_tokens,
            "cost": cost
        }
        self.costs[agent_id]["input_tokens"] += input_tokens
        self.costs[agent_id]["output_tokens"] += output_tokens
        self.costs[agent_id]["cost"] += cost
        self.costs[agent_id]["calls"] += 1

        with open(self.log_file, "a") as f:
            f.write(json.dumps(entry) + "\n")

    def get_daily_cost(self, agent_id=None):
        today = datetime.now().date()
        total = 0.0
        with open(self.log_file) as f:
            for line in f:
                entry = json.loads(line)
                entry_date = datetime.fromisoformat(entry["timestamp"]).date()
                if entry_date == today:
                    if agent_id is None or entry["agent_id"] == agent_id:
                        total += entry["cost"]
        return total

    def get_monthly_cost(self, agent_id=None):
        now = datetime.now()
        total = 0.0
        with open(self.log_file) as f:
            for line in f:
                entry = json.loads(line)
                entry_date = datetime.fromisoformat(entry["timestamp"])
                if entry_date.year == now.year and entry_date.month == now.month:
                    if agent_id is None or entry["agent_id"] == agent_id:
                        total += entry["cost"]
        return total

    def summary(self):
        return {agent: dict(stats) for agent, stats in self.costs.items()}

Budgets and alerts

class BudgetGuard:
    def __init__(self, tracker, daily_budget=10.0, monthly_budget=200.0):
        self.tracker = tracker
        self.daily_budget = daily_budget
        self.monthly_budget = monthly_budget

    def check_budget(self, agent_id=None):
        daily = self.tracker.get_daily_cost(agent_id)
        monthly = self.tracker.get_monthly_cost(agent_id)

        if daily > self.daily_budget:
            self.send_alert("daily", daily, self.daily_budget, agent_id)
            return False

        if monthly > self.monthly_budget:
            self.send_alert("monthly", monthly, self.monthly_budget, agent_id)
            return False

        # Warning at 80% utilization
        if daily > self.daily_budget * 0.8:
            self.send_warning("daily", daily, self.daily_budget, agent_id)
        if monthly > self.monthly_budget * 0.8:
            self.send_warning("monthly", monthly, self.monthly_budget, agent_id)

        return True

    def send_alert(self, budget_type, actual, budget, agent_id):
        print(f"ALERT: {budget_type} budget exceeded!")
        print(f"Agent: {agent_id or 'all'}")
        print(f"Actual: {actual:.2f} $, Budget: {budget:.2f} $")
        # In practice: email, Slack, webhook

    def send_warning(self, budget_type, actual, budget, agent_id):
        print(f"WARNING: {budget_type} budget at 80%!")
        print(f"Agent: {agent_id or 'all'}")
        print(f"Actual: {actual:.2f} $, Budget: {budget:.2f} $")

Agent with Cost Control

class CostControlledAgent:
    def __init__(self, model="gpt-4o", tracker=None, budget_guard=None):
        self.model = model
        self.tracker = tracker or CostTracker()
        self.budget_guard = budget_guard or BudgetGuard(self.tracker)

    def call_model(self, messages, agent_id="default"):
        # Check budget
        if not self.budget_guard.check_budget(agent_id):
            raise BudgetExceededError("Budget exceeded, call rejected")

        # Call model
        result = call_model_with_tracking(messages)
        result["agent_id"] = agent_id

        # Log costs
        self.tracker.log_call(
            agent_id=agent_id,
            model=self.model,
            input_tokens=result["input_tokens"],
            output_tokens=result["output_tokens"],
            cost=result["cost"]
        )

        return result["content"]

class BudgetExceededError(Exception):
    pass

Practical Example 1: Email Agent with Budget

tracker = CostTracker()
guard = BudgetGuard(tracker, daily_budget=5.0, monthly_budget=100.0)
agent = CostControlledAgent(model="gpt-4o-mini", tracker=tracker, budget_guard=guard)

for email in emails:
    try:
        result = agent.call_model(
            [{"role": "user", "content": f"Classify: {email['subject']}"}],
            agent_id="email_classifier"
        )
    except BudgetExceededError:
        print("Budget exceeded, switching to local model")
        # Fallback to Ollama
        result = call_ollama_with_tracking([
            {"role": "user", "content": f"Classify: {email['subject']}"}
        ])

Practical Example 2: Multi-Agent with Distributed Budget

# Different budgets per agent
budgets = {
    "email_classifier": BudgetGuard(tracker, daily_budget=2.0, monthly_budget=50.0),
    "research_agent": BudgetGuard(tracker, daily_budget=10.0, monthly_budget=200.0),
    "writer_agent": BudgetGuard(tracker, daily_budget=5.0, monthly_budget=100.0)
}

agents = {
    name: CostControlledAgent(model="gpt-4o-mini", tracker=tracker, budget_guard=guard)
    for name, guard in budgets.items()
}

Optimizing Costs

1. Use smaller models for simple tasks

# For classification: gpt-4o-mini instead of gpt-4o
# Savings: 0.15 $/M instead of 5 $/M input = 97% cheaper

2. Limit context length

# Don't send entire email threads, just the latest message
messages = messages[-3:]  # Only last 3 messages

3. Caching

from functools import lru_cache

@lru_cache(maxsize=1000)
def cached_classification(email_hash):
    # Don't classify the same email twice
    return call_model_with_tracking([{"role": "user", "content": email_hash}])

4. Local AI for standard tasks

# Standard tasks with Ollama (free)
# Complex tasks with cloud API (paid)

def smart_router(task_complexity):
    if task_complexity == "simple":
        return call_ollama_with_tracking  # Free
    else:
        return call_model_with_tracking  # Cloud

See Local AI vs. API Costs for details.

5. Batch processing

# Bundle multiple requests in a single call
batch_prompt = "Classify the following emails:\n" + "\n".join(
    f"{i+1}. {email}" for i, email in enumerate(emails)
)
# One call instead of 100

Common Pitfalls

  • No tracking: Without measurement, there’s no control. Implement tracking from day one.
  • Wrong model: Using GPT-4o for simple classification is wasteful.
  • No budgets: Without limits, a buggy agent can cost thousands of dollars.
  • No alerts: Without warnings, you won’t notice overspending until the bill arrives.
  • Ignoring context length: Longer contexts cost more. Limit them where possible.
  • No caching: Running the same request multiple times is wasteful.

Further Reading

Key Takeaways:

  • Cost control measures token consumption and API expenses.
  • Budgets and alerts prevent cost explosions.
  • Small models for simple tasks save up to 97%.
  • Caching prevents duplicate calls.
  • Local AI (Ollama) eliminates API costs entirely.

FAQ

Why do I need cost control for AI agents?

AI agents can be expensive, especially in 24/7 operation. Without control, you won’t notice costs until the bill arrives. Cost control measures consumption, sets budgets, and alerts you to overages.

How do I measure token consumption?

Cloud APIs return token counts in the response (usage.prompt_tokens, usage.completion_tokens). Ollama also returns token counts (prompt_eval_count, eval_count).

How do I set up budgets?

Define a daily and monthly budget. Before each model call, check whether the budget is still available. If exceeded, reject the call or switch to a local model.

Does local AI eliminate all costs?

API costs, yes. You only pay for electricity and hardware depreciation. For high-volume usage, local AI is significantly cheaper than cloud APIs.

How do I optimize costs?

Use smaller models for simple tasks, limit context length, cache recurring requests, batch process multiple queries, and use local AI for standard tasks.

How do I set up alerts?

Check on each call whether the budget is exceeded. Send a warning when overspending occurs (email, Slack, webhook). Also alert at 80% capacity to intervene early.

Which model is cheapest?

Local models (Ollama) are free (electricity only). For cloud APIs, GPT-4o-mini (0.15 $/M input) is significantly cheaper than GPT-4o (5 $/M input).

How does caching help?

Caching stores results from model calls. Identical requests aren’t executed twice, saving tokens and costs.

How do I control multiple agents?

Assign each agent its own budget. Track costs per agent and check before each call whether the agent’s budget is still available.

What do I do when the budget is exceeded?

Reject the cloud call and switch to a local model (Ollama). Or pause the agent until the budget becomes available again in the next period.

References and Further Reading

Back to Blog
Share:

Related Posts