Cost Control for AI Agents
What this article covers
- How to monitor token consumption and API costs for AI agents.
- How to set up budgets, alerts, and limits.
- How local AI eliminates API costs.
- Practical examples for cost tracking, budget warnings, and optimization.
- Best practices for cost-efficient agent operation.
Introduction: Understanding cost control for AI agents
AI agents can be expensive. Every agent step calls the language model, every call consumes tokens, and every token costs money when using cloud APIs. An agent running 24/7 that consumes 5 million tokens per day can cost thousands of euros per month. Cost control is the practice of monitoring, limiting, and optimizing this consumption.
This article is for developers who build AI agents and need to manage costs. You should understand what AI agents are and how to run them locally with Ollama. For Python fundamentals, see IRC-Coding.de.
Why do you need cost control?
Imagine you build an agent that sorts emails. It runs 24/7 and calls the model for each email. With 100 emails per day and 10,000 tokens per call, that’s 1 million tokens daily. Using GPT-4o (5 $/M input, 15 $/M output), you’re looking at $20 per day, $600 per month. Without cost control, you won’t notice until the bill arrives.
Cost control for AI agents explained
Cost control means monitoring and limiting the token consumption of AI agents. You measure how many tokens each agent consumes, set budgets, configure alerts, and optimize to reduce spending. With local AI (Ollama), API costs disappear entirely.
The core principle: what you don’t measure, you can’t control.
Who this article is for
- Developers building AI agents who need to manage costs.
- Teams running agents in production and needing to stick to budgets.
- Decision makers who want to understand API costs.
- Self-hosters evaluating local AI as a cost alternative.
Prior experience with AI agents and Ollama is helpful.
Key terms
- Token - A unit of text, roughly 4 characters. Useful to know: it’s the billing unit for APIs.
- AI agents - Programs that call models. Useful to know: what generates costs.
- Ollama - Local model server. Useful to know: eliminates API costs.
- Budget - Cost limit per time period. Useful to know: caps your spending.
- Alert - Warning when budget is exceeded. Useful to know: lets you react early.
- Quantization - Reducing model size. Useful to know: enables faster, cheaper inference.
- Context length - Maximum tokens per call. Useful to know: affects cost.
- Logging - Recording of calls. Useful to know: the foundation for cost tracking.
Measuring token consumption
Cloud API: tokens from the response
Cloud APIs return token counts in their response:
import requests
def call_model_with_tracking(messages):
response = requests.post(
"https://api.openai.com/v1/chat/completions",
headers={"Authorization": "Bearer sk-..."},
json={
"model": "gpt-4o",
"messages": messages,
"stream": False
}
)
data = response.json()
return {
"content": data["choices"][0]["message"]["content"],
"input_tokens": data["usage"]["prompt_tokens"],
"output_tokens": data["usage"]["completion_tokens"],
"cost": calculate_cost("gpt-4o", data["usage"])
}
def calculate_cost(model, usage):
prices = {
"gpt-4o": {"input": 5.0, "output": 15.0}, # per 1M tokens
"gpt-4o-mini": {"input": 0.15, "output": 0.60},
"claude-3-5-sonnet": {"input": 3.0, "output": 15.0}
}
p = prices.get(model, {"input": 0, "output": 0})
return (usage["prompt_tokens"] * p["input"] + usage["completion_tokens"] * p["output"]) / 1_000_000
Ollama: tokens from the response
Ollama also returns token counts:
def call_ollama_with_tracking(messages):
response = requests.post(
"http://localhost:11434/api/chat",
json={"model": "llama3.1", "messages": messages, "stream": False}
)
data = response.json()
return {
"content": data["message"]["content"],
"input_tokens": data.get("prompt_eval_count", 0),
"output_tokens": data.get("eval_count", 0),
"cost": 0.0 # Local AI: no API costs
}
Implementing a cost tracker
from datetime import datetime, timedelta
from collections import defaultdict
import json
class CostTracker:
def __init__(self, log_file="agent_costs.json"):
self.log_file = log_file
self.costs = defaultdict(lambda: {"input_tokens": 0, "output_tokens": 0, "cost": 0.0, "calls": 0})
def log_call(self, agent_id, model, input_tokens, output_tokens, cost):
entry = {
"timestamp": datetime.now().isoformat(),
"agent_id": agent_id,
"model": model,
"input_tokens": input_tokens,
"output_tokens": output_tokens,
"cost": cost
}
self.costs[agent_id]["input_tokens"] += input_tokens
self.costs[agent_id]["output_tokens"] += output_tokens
self.costs[agent_id]["cost"] += cost
self.costs[agent_id]["calls"] += 1
with open(self.log_file, "a") as f:
f.write(json.dumps(entry) + "\n")
def get_daily_cost(self, agent_id=None):
today = datetime.now().date()
total = 0.0
with open(self.log_file) as f:
for line in f:
entry = json.loads(line)
entry_date = datetime.fromisoformat(entry["timestamp"]).date()
if entry_date == today:
if agent_id is None or entry["agent_id"] == agent_id:
total += entry["cost"]
return total
def get_monthly_cost(self, agent_id=None):
now = datetime.now()
total = 0.0
with open(self.log_file) as f:
for line in f:
entry = json.loads(line)
entry_date = datetime.fromisoformat(entry["timestamp"])
if entry_date.year == now.year and entry_date.month == now.month:
if agent_id is None or entry["agent_id"] == agent_id:
total += entry["cost"]
return total
def summary(self):
return {agent: dict(stats) for agent, stats in self.costs.items()}
Budgets and alerts
class BudgetGuard:
def __init__(self, tracker, daily_budget=10.0, monthly_budget=200.0):
self.tracker = tracker
self.daily_budget = daily_budget
self.monthly_budget = monthly_budget
def check_budget(self, agent_id=None):
daily = self.tracker.get_daily_cost(agent_id)
monthly = self.tracker.get_monthly_cost(agent_id)
if daily > self.daily_budget:
self.send_alert("daily", daily, self.daily_budget, agent_id)
return False
if monthly > self.monthly_budget:
self.send_alert("monthly", monthly, self.monthly_budget, agent_id)
return False
# Warning at 80% utilization
if daily > self.daily_budget * 0.8:
self.send_warning("daily", daily, self.daily_budget, agent_id)
if monthly > self.monthly_budget * 0.8:
self.send_warning("monthly", monthly, self.monthly_budget, agent_id)
return True
def send_alert(self, budget_type, actual, budget, agent_id):
print(f"ALERT: {budget_type} budget exceeded!")
print(f"Agent: {agent_id or 'all'}")
print(f"Actual: {actual:.2f} $, Budget: {budget:.2f} $")
# In practice: email, Slack, webhook
def send_warning(self, budget_type, actual, budget, agent_id):
print(f"WARNING: {budget_type} budget at 80%!")
print(f"Agent: {agent_id or 'all'}")
print(f"Actual: {actual:.2f} $, Budget: {budget:.2f} $")
Agent with Cost Control
class CostControlledAgent:
def __init__(self, model="gpt-4o", tracker=None, budget_guard=None):
self.model = model
self.tracker = tracker or CostTracker()
self.budget_guard = budget_guard or BudgetGuard(self.tracker)
def call_model(self, messages, agent_id="default"):
# Check budget
if not self.budget_guard.check_budget(agent_id):
raise BudgetExceededError("Budget exceeded, call rejected")
# Call model
result = call_model_with_tracking(messages)
result["agent_id"] = agent_id
# Log costs
self.tracker.log_call(
agent_id=agent_id,
model=self.model,
input_tokens=result["input_tokens"],
output_tokens=result["output_tokens"],
cost=result["cost"]
)
return result["content"]
class BudgetExceededError(Exception):
pass
Practical Example 1: Email Agent with Budget
tracker = CostTracker()
guard = BudgetGuard(tracker, daily_budget=5.0, monthly_budget=100.0)
agent = CostControlledAgent(model="gpt-4o-mini", tracker=tracker, budget_guard=guard)
for email in emails:
try:
result = agent.call_model(
[{"role": "user", "content": f"Classify: {email['subject']}"}],
agent_id="email_classifier"
)
except BudgetExceededError:
print("Budget exceeded, switching to local model")
# Fallback to Ollama
result = call_ollama_with_tracking([
{"role": "user", "content": f"Classify: {email['subject']}"}
])
Practical Example 2: Multi-Agent with Distributed Budget
# Different budgets per agent
budgets = {
"email_classifier": BudgetGuard(tracker, daily_budget=2.0, monthly_budget=50.0),
"research_agent": BudgetGuard(tracker, daily_budget=10.0, monthly_budget=200.0),
"writer_agent": BudgetGuard(tracker, daily_budget=5.0, monthly_budget=100.0)
}
agents = {
name: CostControlledAgent(model="gpt-4o-mini", tracker=tracker, budget_guard=guard)
for name, guard in budgets.items()
}
Optimizing Costs
1. Use smaller models for simple tasks
# For classification: gpt-4o-mini instead of gpt-4o
# Savings: 0.15 $/M instead of 5 $/M input = 97% cheaper
2. Limit context length
# Don't send entire email threads, just the latest message
messages = messages[-3:] # Only last 3 messages
3. Caching
from functools import lru_cache
@lru_cache(maxsize=1000)
def cached_classification(email_hash):
# Don't classify the same email twice
return call_model_with_tracking([{"role": "user", "content": email_hash}])
4. Local AI for standard tasks
# Standard tasks with Ollama (free)
# Complex tasks with cloud API (paid)
def smart_router(task_complexity):
if task_complexity == "simple":
return call_ollama_with_tracking # Free
else:
return call_model_with_tracking # Cloud
See Local AI vs. API Costs for details.
5. Batch processing
# Bundle multiple requests in a single call
batch_prompt = "Classify the following emails:\n" + "\n".join(
f"{i+1}. {email}" for i, email in enumerate(emails)
)
# One call instead of 100
Common Pitfalls
- No tracking: Without measurement, there’s no control. Implement tracking from day one.
- Wrong model: Using GPT-4o for simple classification is wasteful.
- No budgets: Without limits, a buggy agent can cost thousands of dollars.
- No alerts: Without warnings, you won’t notice overspending until the bill arrives.
- Ignoring context length: Longer contexts cost more. Limit them where possible.
- No caching: Running the same request multiple times is wasteful.
Further Reading
- AI Agents Fundamentals - What AI agents are.
- AI Agents with Ollama - Local AI without API costs.
- Logging - Foundation for cost tracking.
- Local AI vs. API Costs - Cost comparison.
- Quantization - Reduce model size.
- Context Length - Understanding context.
- Cloud Cost Calculator - Calculate API costs.
- Electricity Cost Calculator - Calculate power costs for local AI.
Key Takeaways:
- Cost control measures token consumption and API expenses.
- Budgets and alerts prevent cost explosions.
- Small models for simple tasks save up to 97%.
- Caching prevents duplicate calls.
- Local AI (Ollama) eliminates API costs entirely.
FAQ
Why do I need cost control for AI agents?
How do I measure token consumption?
How do I set up budgets?
Does local AI eliminate all costs?
How do I optimize costs?
How do I set up alerts?
Which model is cheapest?
How does caching help?
How do I control multiple agents?
What do I do when the budget is exceeded?
References and Further Reading
- OpenAI Pricing - API pricing.
- Ollama - Local model server.
- Cloud Cost Calculator - Calculate API costs.
- Local AI vs. API Costs - Cost comparison.


