Using Ollama for AI Agents
What this article covers
- How Ollama works as a model server for AI agents.
- Which frameworks support Ollama and how to connect them.
- How Function Calling works with Ollama and which models are suitable.
- How to build a simple agent with Ollama and LangGraph.
- Common pitfalls when working with local agents and how to solve them.
Introduction: Ollama as a model server for AI agents
AI agents are programs that plan independently, invoke tools, and complete tasks. They need a language model that makes decisions and generates tool calls. Ollama is a local model server that does exactly that: it provides a model accessible via REST API and supports Function Calling.
This article is for developers building AI agents with local models. You should understand what AI agents are, how tool calling works, and how to install Ollama. Basic Python knowledge is helpful for the code examples. If you’re new to Python, check out tutorials on IRC-Coding.de.
Why use Ollama for AI agents?
Imagine building an agent that reads, classifies, and routes emails to the right department. The agent uses a language model to determine the category and calls functions to forward the email.
If you rent the model from a cloud provider, you pay per request. With 1000 emails per day at 0.001 EUR per request, that’s 1 EUR per day and 365 EUR per year. Add data protection concerns: every email goes to an external provider.
With Ollama, you run the model locally. After the hardware investment, you only pay for electricity. Emails never leave your network. You’re independent from API limits, price increases, or service shutdowns.
Ollama for AI agents explained
Ollama is a model server running locally on your hardware. It provides a REST API through which you can query models. For AI agents, Function Calling is particularly important: the model can define functions it’s allowed to call and generates structured invocations that your code executes.
The core idea is simple: Ollama provides the model, your code provides the logic, and the model decides which tools to invoke.
Who should use Ollama for AI agents?
- Developers building AI agents with local models.
- Enterprises that can’t move privacy-critical workloads to the cloud.
- Hobbyists building agents for personal automation.
- Researchers evaluating models for agent workflows.
Prior experience with Python, APIs, and basic prompt engineering is helpful. If you haven’t worked with Function Calling before, read the fundamentals first.
Key terms around Ollama and AI agents
- Ollama - Local model server. Useful as: the model backend for any local agent.
- Function Calling - Model’s ability to invoke functions. Useful when: the agent needs tools.
- LangGraph - Graph-based agent framework. Useful for: complex agents with state management.
- CrewAI - Multi-agent framework. Useful for: agents working together.
- OpenClaw - Open-source agent platform. Useful for: self-hosting agents.
- MCP - Model Context Protocol. Useful for: connecting agents with external tools.
- Ollama REST API - HTTP API for model requests. Useful when: calling Ollama directly.
Which frameworks support Ollama?
Most popular agent frameworks support Ollama as a model backend. Integration is usually straightforward because Ollama provides an OpenAI-compatible API.
| Framework | Ollama integration | Use case |
|---|---|---|
| LangGraph | Direct, via ChatOllama | Complex agents with state |
| CrewAI | Via LLM class | Multi-agent systems |
| LangChain | Via ChatOllama | General AI applications |
| LlamaIndex | Via Ollama class | RAG applications |
| OpenClaw | Configurable | Self-hosted agents |
| OpenHands | Configurable | Software development |
Integration typically works like this: install the framework, configure Ollama as your model backend, and define your tools. The framework calls Ollama when the model needs to make a decision and executes the generated tool calls.
Function Calling with Ollama
Function Calling is the most important capability for AI agents. The model receives a list of functions it can invoke and decides based on the user query which function is relevant. It generates a structured call that your code executes.
Ollama has supported Function Calling since version 0.3.0. Models well-suited for this include:
- Llama 3.1 8B - Small, fast, good for simple agents.
- Qwen 2.5 14B - Larger, better at complex tool invocations.
- Mistral Nemo 12B - Balanced, strong at tool selection.
- Llama 3.3 70B - Large, if your hardware permits.
Here’s an example using the Ollama REST API:
import requests
import json
OLLAMA_URL = "http://localhost:11434/api/chat"
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Holt das aktuelle Wetter für einen Ort",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "Ortsname, z.B. 'Berlin'"
}
},
"required": ["location"]
}
}
}
]
response = requests.post(OLLAMA_URL, json={
"model": "llama3.1",
"messages": [
{"role": "user", "content": "Wie ist das Wetter in Berlin?"}
],
"tools": tools,
"stream": False
})
result = response.json()
print(json.dumps(result, indent=2, ensure_ascii=False))
The model generates a tool call included in the response. Your code executes the get_weather function and sends the result back to the model, which uses it to generate the final response.
Practical example: An agent with Ollama and LangGraph
Imagine you want an agent that conducts research. It searches the internet, summarizes results, and generates a final answer. With Ollama and LangGraph, it looks like this:
from langchain_ollama import ChatOllama
from langgraph.graph import StateGraph, END
from typing import TypedDict, Annotated
import operator
class AgentState(TypedDict):
messages: Annotated[list, operator.add]
next_step: str
llm = ChatOllama(model="llama3.1", temperature=0)
def search_node(state: AgentState):
# Simulierte Suchfunktion
query = state["messages"][-1]["content"]
results = f"Suchergebnisse für: {query}"
return {"messages": [{"role": "system", "content": results}]}
def summarize_node(state: AgentState):
# Zusammenfassung mit LLM
context = "\n".join([m["content"] for m in state["messages"]])
response = llm.invoke(f"Fasse zusammen: {context}")
return {"messages": [{"role": "assistant", "content": response.content}]}
def should_continue(state: AgentState):
if "suche" in state["messages"][-1]["content"].lower():
return "search"
return "summarize"
workflow = StateGraph(AgentState)
workflow.add_node("search", search_node)
workflow.add_node("summarize", summarize_node)
workflow.set_entry_point("search")
workflow.add_conditional_edges("search", should_continue)
workflow.add_edge("summarize", END)
app = workflow.compile()
result = app.invoke({"messages": [{"role": "user", "content": "Suche nach lokaler KI"}]})
print(result["messages"][-1]["content"])
This agent cycles through multiple states: search, summarize, finish. Ollama provides the model, LangGraph orchestrates the workflow. The model decides whether additional steps are needed.
Connecting Ollama with MCP
The Model Context Protocol (MCP) is a standard that enables agents to access external tools. Instead of writing tool definitions for each agent individually, you use MCP servers that provide tools.
Ollama itself isn’t an MCP client, but most agent frameworks that support Ollama can also use MCP. LangGraph, CrewAI, and OpenClaw have MCP integrations.
A typical setup looks like this:
- Ollama runs as a model server on
localhost:11434. - MCP servers provide tools (for example, file system, database, web search).
- Agent framework connects Ollama with the MCP server.
- The agent calls Ollama when it needs to make a decision.
- The model generates a tool call, which the framework executes through MCP.
For details on building your own MCP server, see the article Building MCP Servers.
Choosing Models for Agents
Model selection has a massive impact on agent quality. A model that’s too small won’t understand tool calls properly, while one that’s too large will be slow and consume lots of VRAM.
| Model | Parameters | VRAM (Q4) | Best for |
|---|---|---|---|
| Llama 3.1 8B | 8B | 6 GB | Simple agents, classification |
| Qwen 2.5 14B | 14B | 10 GB | Medium agents, tool selection |
| Mistral Nemo 12B | 12B | 8 GB | Balanced agents |
| Llama 3.3 70B | 70B | 40 GB | Complex agents, if hardware permits |
| Qwen 2.5 32B | 32B | 20 GB | Demanding tool calls |
Always test your model before deploying it. The article Ollama Evaluation shows how to systematically compare models. For function calling, test your model with your specific tools, because not every model understands every tool schema equally well.
Common Pitfalls with Ollama Agents
- Wrong model choice: A 7B model won’t handle complex tool calls well. Test whether your model understands your tools correctly before deployment.
- Missing error handling: When a tool call fails, the agent must detect it and respond. An agent that crashes on every error isn’t production-ready.
- Too many tools: A model with 20 tools becomes overwhelmed. Prefer a few well-defined tools over many vague ones.
- No conversation limits: An agent that keeps thinking indefinitely burns through tokens and time. Set a maximum step count.
- Ignoring security: An agent that can call tools can also cause damage. Tool permissions and sandboxing are mandatory.
- No monitoring: If you don’t notice when an agent makes wrong decisions, you’ll only find out after harm is done. Audit logging matters.
- Forgetting human approval: Critical actions need human approval, not blind execution.
Further Reading and Resources on Ollama Agents
- Ollama Function Calling - Defining tools in Ollama.
- Ollama REST API - API for model requests.
- LangGraph Framework - Building complex agents.
- CrewAI Framework - Multi-agent systems.
- MCP Basics - Providing tools for agents.
- Ollama Evaluation - Systematically comparing models.
- Agent Security - Securing agents.
Key Takeaways:
- Ollama is a local model server that supports function calling for AI agents.
- Most popular frameworks (LangGraph, CrewAI, LangChain) support Ollama.
- Model choice heavily influences agent quality.
- MCP connects agents to external tools without requiring you to define each tool yourself.
- Security, error handling, and human approval are essential for agents with tool access.
FAQ: Ollama for AI Agents - Common Questions
What is Ollama for AI agents?
Which frameworks support Ollama?
Does Ollama support function calling?
Which model is suitable for AI agents?
Can I use Ollama with MCP?
What hardware do I need for Ollama agents?
How do I secure Ollama agents?
What does it cost to run Ollama agents?
How do I handle errors in agents?
Can I run multiple agents with Ollama?
References and Further Reading
- Ollama Documentation - Official Ollama docs.
- LangGraph - Agent framework.
- CrewAI - Multi-agent framework.
- Model Context Protocol - Tool standard for agents.
- LangChain Ollama Integration - Ollama in LangChain.


