Running AI Agents Locally with Ollama
What This Article Covers
- How to run AI agents locally using Ollama as your backend.
- Which agent frameworks work with Ollama and how to set them up.
- Implementing tool use, memory, and multi-agent systems.
- Practical examples for autonomous agents: research, code execution, file management.
- Best practices for security, performance, and reliability.
Introduction: Understanding AI Agents with Ollama
AI agents are programs that leverage a language model to solve tasks independently. They plan, execute tools, observe results, and adjust their approach. With Ollama as a backend, these agents run entirely locally, no cloud, no API costs, no data privacy concerns.
This article is for developers who want to build autonomous agents using local AI. You should be familiar with AI agent basics, how Ollama works, and what Function Calling is. For Python fundamentals, check out IRC-Coding.de.
Why Use Local AI Agents?
Imagine building an agent that sorts your emails, manages appointments, and summarizes documents. In the cloud, every agent step costs money, every API call is logged, every data point flows to a provider. With local AI, everything runs on your hardware, no per-call charges, no data leaving your system.
Local agents shine when you handle sensitive data, need many agent steps (expensive in the cloud), or want independence from cloud providers.
How AI Agents Work with Ollama
An AI agent executes a loop: plan → call tool → observe result → adjust plan. The language model makes decisions; code executes the tools. With Ollama as the backend, the language model runs locally, and tools are local programs (filesystem, shell, web requests).
The core concept: the agent is the loop, Ollama is the brain, tools are the hands.
Who Should Read This?
- Developers building autonomous agents with local AI.
- System administrators running agents for automation on their own hardware.
- Researchers investigating agent behavior using local models.
- Self-hosters seeking independence from cloud providers.
Prior knowledge of Python, Ollama, and Function Calling is assumed.
Key Concepts for AI Agents with Ollama
- AI Agent - A program that solves tasks autonomously. Useful for: autonomous workflows.
- Ollama - Local model server. Useful for: the agent’s brain.
- Function Calling - Structured AI responses. Useful for: enabling agents to invoke tools.
- Tool - A function the agent can call. Useful for: the agent’s hands.
- Memory - The agent’s memory. Useful for: recalling prior steps.
- MCP - Model Context Protocol. Useful for: a standard for tool integration.
- OpenHands - Open-source agent platform. Useful for: a ready-made agent with Ollama support.
- Agent Loop - The core loop: plan → tool → observe → adapt. Useful for: the heart of any agent.
Agent Frameworks with Ollama Support
LangChain with Ollama
LangChain is the most popular framework for LLM applications. It supports Ollama as a backend.
from langchain_community.llms import Ollama
from langchain.agents import AgentExecutor, create_react_agent
from langchain.tools import Tool
# Ollama as backend
llm = Ollama(model="llama3.1", base_url="http://localhost:11434")
# Define tools
def search_wikipedia(query):
# Implementation here
return f"Result for: {query}"
tools = [
Tool(name="Wikipedia", func=search_wikipedia, description="Search Wikipedia")
]
# Create agent
agent = create_react_agent(llm, tools, prompt)
executor = AgentExecutor(agent=agent, tools=tools)
# Run agent
result = executor.invoke({"input": "What is the capital of France?"})
print(result)
LangGraph with Ollama
LangGraph extends LangChain for stateful agents. It allows complex agent graphs with cycles.
from langchain_community.llms import Ollama
from langgraph.graph import StateGraph, END
llm = Ollama(model="llama3.1")
# Define state
from typing import TypedDict, List
class AgentState(TypedDict):
messages: List[str]
tool_results: List[str]
# Define nodes
def call_model(state):
response = llm.invoke(state["messages"][-1])
return {"messages": state["messages"] + [response]}
def call_tool(state):
# Tool execution
return {"tool_results": ["Result"]}
# Build graph
workflow = StateGraph(AgentState)
workflow.add_node("model", call_model)
workflow.add_node("tool", call_tool)
workflow.add_edge("model", "tool")
workflow.add_edge("tool", "model")
CrewAI with Ollama
CrewAI is a framework for multi-agent systems. Multiple agents with different roles collaborate.
from crewai import Agent, Task, Crew
from crewai.llms import Ollama
# Configure LLM
llm = Ollama(model="llama3.1", base_url="http://localhost:11434")
# Define agents
researcher = Agent(
role="Researcher",
goal="Gather information",
backstory="You are an experienced researcher.",
llm=llm
)
writer = Agent(
role="Writer",
goal="Write a report",
backstory="You are an experienced writer.",
llm=llm
)
# Define tasks
research_task = Task(description="Research topic X", agent=researcher)
write_task = Task(description="Write a report", agent=writer)
# Create crew
crew = Crew(agents=[researcher, writer], tasks=[research_task, write_task])
result = crew.kickoff()
OpenHands with Ollama
OpenHands (formerly OpenDevin) is a complete agent platform with a web UI. It supports Ollama as a backend.
See Setting up OpenHands for details.
Tool Use with Ollama
Function Calling with Ollama
Ollama supports Function Calling, which agents use to invoke tools.
import requests
def call_ollama_with_tools(messages, tools):
response = requests.post(
"http://localhost:11434/api/chat",
json={
"model": "llama3.1",
"messages": messages,
"tools": tools,
"stream": False
}
)
return response.json()
# Define tools
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"}
},
"required": ["city"]
}
}
}
]
# Agent loop
messages = [{"role": "user", "content": "What is the weather in Berlin?"}]
while True:
response = call_ollama_with_tools(messages, tools)
msg = response["message"]
if msg.get("tool_calls"):
for tool_call in msg["tool_calls"]:
# Execute tool
result = execute_tool(tool_call)
messages.append(msg)
messages.append({
"role": "tool",
"content": result
})
else:
print(msg["content"])
break
MCP with Ollama
The Model Context Protocol (MCP) is a standard for tool integration. Ollama models can leverage MCP servers.
See MCP Basics for details.
Memory for Agents
Agents need memory to recall earlier steps. Several memory patterns exist:
Short-Term Memory
The current context, including the most recent messages. Limited by the model’s context window.
messages = [
{"role": "system", "content": "Du bist ein Agent."},
{"role": "user", "content": "Schritt 1"},
{"role": "assistant", "content": "Ergebnis 1"},
{"role": "user", "content": "Schritt 2"},
# ...
]
Long-Term Memory
Persistent storage, such as in a vector database. The agent can retrieve earlier insights and findings.
from chromadb import Client
chroma = Client()
collection = chroma.create_collection("agent_memory")
# Store a memory
collection.add(
documents=["Wichtige Erkenntnis aus Schritt 5"],
ids=["step_5"]
)
# Retrieve a memory
results = collection.query(query_texts=["Was weiß ich über X?"], n_results=3)
Summary Memory
When context grows too long, the agent summarizes older messages.
def summarize_if_needed(messages, max_length=4000):
total_length = sum(len(m["content"]) for m in messages)
if total_length > max_length:
# Summarize
summary = call_ollama([
{"role": "system", "content": "Fasse diese Konversation zusammen."},
{"role": "user", "content": str(messages)}
])
messages = [
messages[0], # Keep system prompt
{"role": "system", "content": f"Zusammenfassung: {summary}"}
]
return messages
Multi-Agent Setups
Hierarchical Agents
A supervisor agent delegates work to specialized agents:
def supervisor(task):
# Supervisor decides which agent handles the task
decision = call_ollama([
{"role": "system", "content": "Du bist ein Supervisor. Entscheide, welcher Agent dran ist: researcher, coder, oder writer."},
{"role": "user", "content": task}
])
if "researcher" in decision:
return researcher_agent(task)
elif "coder" in decision:
return coder_agent(task)
elif "writer" in decision:
return writer_agent(task)
Parallel Agents
Multiple agents work in parallel on the same problem:
import concurrent.futures
def parallel_agents(task, agents):
with concurrent.futures.ThreadPoolExecutor() as executor:
futures = [executor.submit(agent, task) for agent in agents]
results = [f.result() for f in futures]
return results
Practical Example 1: Research Agent
An agent that researches a topic, collects sources, and produces a report.
def research_agent(topic):
messages = [
{"role": "system", "content": "Du bist ein Recherche-Agent. Nutze die Tools search_web und read_page."},
{"role": "user", "content": f"Recherchiere: {topic}"}
]
tools = [
{"type": "function", "function": {"name": "search_web", "parameters": {"query": {"type": "string"}}}},
{"type": "function", "function": {"name": "read_page", "parameters": {"url": {"type": "string"}}}}
]
while True:
response = call_ollama_with_tools(messages, tools)
msg = response["message"]
messages.append(msg)
if msg.get("tool_calls"):
for tc in msg["tool_calls"]:
result = execute_tool(tc)
messages.append({"role": "tool", "content": result})
else:
return msg["content"]
Practical Example 2: Code Execution Agent
An agent that writes and runs code to solve tasks.
def coder_agent(task):
messages = [
{"role": "system", "content": "Du bist ein Code-Agent. Nutze write_code und run_code."},
{"role": "user", "content": task}
]
tools = [
{"type": "function", "function": {"name": "write_code", "parameters": {"code": {"type": "string"}}}},
{"type": "function", "function": {"name": "run_code", "parameters": {"file": {"type": "string"}}}}
]
# Agent loop as above
# IMPORTANT: Run code in a sandbox!
See Sandbox Environments for security guidance.
Security Considerations
- Sandbox for Code Execution: Agents that execute code must run in a sandbox. See Sandbox Environments.
- Least Privilege: Agents should only receive tools they actually need. See Least Privilege.
- Human-in-the-Loop: For critical actions, require human approval. See Human Approval.
- Audit Logging: Log all agent actions. See Audit Logging.
- Guardrails: Use guardrails to prevent harmful actions. See Configuring Guardrails.
- Prompt Injection Protection: Defend against prompt injection attacks. See Prompt Injection Protection.
Performance Tips
- Small Model for Planning: Use a smaller model for planning steps, a larger one for complex reasoning tasks.
- Watch Context Length: Long agent loops exhaust context windows. Use summary memory.
- Cache Results: Cache tool outputs to avoid redundant calls.
- Parallelize: Execute tools in parallel whenever possible.
- Choose Your Model: For tool use, pick models with strong function calling support, such as
llama3.1,qwen2.5, ormistral.
Common Pitfalls
- No Sandbox: Running code-executing agents without a sandbox is a serious security risk.
- Oversized Models: Using a 70B model for every agent step wastes VRAM and slows execution.
- Missing Error Handling: If a tool fails, the agent should adapt, not crash.
- Exceeding Context Limits: Long loops blow through the context window. Use summary memory.
- Infinite Loops: Agents can get stuck in cycles. Set a maximum step limit.
- No Guardrails: Without guardrails, agents can take harmful actions.
Further Reading and Resources on AI Agents with Ollama
- Install Ollama - Agent runtime.
- Function Calling - Tool-use foundations.
- MCP Basics - Model Context Protocol.
- AI Agent Basics - Agent fundamentals.
- Set Up OpenHands - Ready-made agent platform.
- Sandbox Environments - Securing code agents.
- Human Approval - Human-in-the-loop patterns.
- Configuring Guardrails - Agent guardrails.
Key Takeaways:
- AI agents with Ollama run entirely local, no cloud required.
- Frameworks: LangChain, LangGraph, CrewAI, OpenHands.
- Tool use via function calling or MCP.
- Memory patterns: short-term, long-term, summary.
- Multi-agent architectures: hierarchical or parallel.
- Security essentials: sandboxes, least privilege, guardrails, audit logging.
FAQ: AI Agents with Ollama
What are AI agents?
Can I use Ollama as a backend for agents?
Which frameworks support Ollama?
How do agents call tools?
How does memory work for agents?
Can I run multiple agents in parallel?
How do I secure AI agents?
Which model is right for agents?
What does running local agents cost?
What should I do about infinite loops?
Resources and Further Reading
- Ollama Documentation - Official Ollama docs.
- LangChain Documentation - LangChain docs.
- CrewAI Documentation - CrewAI docs.
- OpenHands - Agent platform.


