Skip to content
BotServBotServ
OllamaAI AgentsFunction CallingLangGraphCrewAI

Using Ollama for AI Agents

Deploy Ollama as a model server for AI agents. Function calling, tool integration, frameworks and practical examples.

S

schutzgeist

9 min read
Using Ollama for AI Agents

Using Ollama for AI Agents

What this article covers

  • How Ollama works as a model server for AI agents.
  • Which frameworks support Ollama and how to connect them.
  • How Function Calling works with Ollama and which models are suitable.
  • How to build a simple agent with Ollama and LangGraph.
  • Common pitfalls when working with local agents and how to solve them.

Introduction: Ollama as a model server for AI agents

AI agents are programs that plan independently, invoke tools, and complete tasks. They need a language model that makes decisions and generates tool calls. Ollama is a local model server that does exactly that: it provides a model accessible via REST API and supports Function Calling.

This article is for developers building AI agents with local models. You should understand what AI agents are, how tool calling works, and how to install Ollama. Basic Python knowledge is helpful for the code examples. If you’re new to Python, check out tutorials on IRC-Coding.de.

Why use Ollama for AI agents?

Imagine building an agent that reads, classifies, and routes emails to the right department. The agent uses a language model to determine the category and calls functions to forward the email.

If you rent the model from a cloud provider, you pay per request. With 1000 emails per day at 0.001 EUR per request, that’s 1 EUR per day and 365 EUR per year. Add data protection concerns: every email goes to an external provider.

With Ollama, you run the model locally. After the hardware investment, you only pay for electricity. Emails never leave your network. You’re independent from API limits, price increases, or service shutdowns.

Ollama for AI agents explained

Ollama is a model server running locally on your hardware. It provides a REST API through which you can query models. For AI agents, Function Calling is particularly important: the model can define functions it’s allowed to call and generates structured invocations that your code executes.

The core idea is simple: Ollama provides the model, your code provides the logic, and the model decides which tools to invoke.

Who should use Ollama for AI agents?

  • Developers building AI agents with local models.
  • Enterprises that can’t move privacy-critical workloads to the cloud.
  • Hobbyists building agents for personal automation.
  • Researchers evaluating models for agent workflows.

Prior experience with Python, APIs, and basic prompt engineering is helpful. If you haven’t worked with Function Calling before, read the fundamentals first.

Key terms around Ollama and AI agents

  • Ollama - Local model server. Useful as: the model backend for any local agent.
  • Function Calling - Model’s ability to invoke functions. Useful when: the agent needs tools.
  • LangGraph - Graph-based agent framework. Useful for: complex agents with state management.
  • CrewAI - Multi-agent framework. Useful for: agents working together.
  • OpenClaw - Open-source agent platform. Useful for: self-hosting agents.
  • MCP - Model Context Protocol. Useful for: connecting agents with external tools.
  • Ollama REST API - HTTP API for model requests. Useful when: calling Ollama directly.

Which frameworks support Ollama?

Most popular agent frameworks support Ollama as a model backend. Integration is usually straightforward because Ollama provides an OpenAI-compatible API.

FrameworkOllama integrationUse case
LangGraphDirect, via ChatOllamaComplex agents with state
CrewAIVia LLM classMulti-agent systems
LangChainVia ChatOllamaGeneral AI applications
LlamaIndexVia Ollama classRAG applications
OpenClawConfigurableSelf-hosted agents
OpenHandsConfigurableSoftware development

Integration typically works like this: install the framework, configure Ollama as your model backend, and define your tools. The framework calls Ollama when the model needs to make a decision and executes the generated tool calls.

Function Calling with Ollama

Function Calling is the most important capability for AI agents. The model receives a list of functions it can invoke and decides based on the user query which function is relevant. It generates a structured call that your code executes.

Ollama has supported Function Calling since version 0.3.0. Models well-suited for this include:

  • Llama 3.1 8B - Small, fast, good for simple agents.
  • Qwen 2.5 14B - Larger, better at complex tool invocations.
  • Mistral Nemo 12B - Balanced, strong at tool selection.
  • Llama 3.3 70B - Large, if your hardware permits.

Here’s an example using the Ollama REST API:

import requests
import json

OLLAMA_URL = "http://localhost:11434/api/chat"

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Holt das aktuelle Wetter für einen Ort",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {
                        "type": "string",
                        "description": "Ortsname, z.B. 'Berlin'"
                    }
                },
                "required": ["location"]
            }
        }
    }
]

response = requests.post(OLLAMA_URL, json={
    "model": "llama3.1",
    "messages": [
        {"role": "user", "content": "Wie ist das Wetter in Berlin?"}
    ],
    "tools": tools,
    "stream": False
})

result = response.json()
print(json.dumps(result, indent=2, ensure_ascii=False))

The model generates a tool call included in the response. Your code executes the get_weather function and sends the result back to the model, which uses it to generate the final response.

Practical example: An agent with Ollama and LangGraph

Imagine you want an agent that conducts research. It searches the internet, summarizes results, and generates a final answer. With Ollama and LangGraph, it looks like this:

from langchain_ollama import ChatOllama
from langgraph.graph import StateGraph, END
from typing import TypedDict, Annotated
import operator

class AgentState(TypedDict):
    messages: Annotated[list, operator.add]
    next_step: str

llm = ChatOllama(model="llama3.1", temperature=0)

def search_node(state: AgentState):
    # Simulierte Suchfunktion
    query = state["messages"][-1]["content"]
    results = f"Suchergebnisse für: {query}"
    return {"messages": [{"role": "system", "content": results}]}

def summarize_node(state: AgentState):
    # Zusammenfassung mit LLM
    context = "\n".join([m["content"] for m in state["messages"]])
    response = llm.invoke(f"Fasse zusammen: {context}")
    return {"messages": [{"role": "assistant", "content": response.content}]}

def should_continue(state: AgentState):
    if "suche" in state["messages"][-1]["content"].lower():
        return "search"
    return "summarize"

workflow = StateGraph(AgentState)
workflow.add_node("search", search_node)
workflow.add_node("summarize", summarize_node)
workflow.set_entry_point("search")
workflow.add_conditional_edges("search", should_continue)
workflow.add_edge("summarize", END)

app = workflow.compile()
result = app.invoke({"messages": [{"role": "user", "content": "Suche nach lokaler KI"}]})
print(result["messages"][-1]["content"])

This agent cycles through multiple states: search, summarize, finish. Ollama provides the model, LangGraph orchestrates the workflow. The model decides whether additional steps are needed.

Connecting Ollama with MCP

The Model Context Protocol (MCP) is a standard that enables agents to access external tools. Instead of writing tool definitions for each agent individually, you use MCP servers that provide tools.

Ollama itself isn’t an MCP client, but most agent frameworks that support Ollama can also use MCP. LangGraph, CrewAI, and OpenClaw have MCP integrations.

A typical setup looks like this:

  1. Ollama runs as a model server on localhost:11434.
  2. MCP servers provide tools (for example, file system, database, web search).
  3. Agent framework connects Ollama with the MCP server.
  4. The agent calls Ollama when it needs to make a decision.
  5. The model generates a tool call, which the framework executes through MCP.

For details on building your own MCP server, see the article Building MCP Servers.

Choosing Models for Agents

Model selection has a massive impact on agent quality. A model that’s too small won’t understand tool calls properly, while one that’s too large will be slow and consume lots of VRAM.

ModelParametersVRAM (Q4)Best for
Llama 3.1 8B8B6 GBSimple agents, classification
Qwen 2.5 14B14B10 GBMedium agents, tool selection
Mistral Nemo 12B12B8 GBBalanced agents
Llama 3.3 70B70B40 GBComplex agents, if hardware permits
Qwen 2.5 32B32B20 GBDemanding tool calls

Always test your model before deploying it. The article Ollama Evaluation shows how to systematically compare models. For function calling, test your model with your specific tools, because not every model understands every tool schema equally well.

Common Pitfalls with Ollama Agents

  • Wrong model choice: A 7B model won’t handle complex tool calls well. Test whether your model understands your tools correctly before deployment.
  • Missing error handling: When a tool call fails, the agent must detect it and respond. An agent that crashes on every error isn’t production-ready.
  • Too many tools: A model with 20 tools becomes overwhelmed. Prefer a few well-defined tools over many vague ones.
  • No conversation limits: An agent that keeps thinking indefinitely burns through tokens and time. Set a maximum step count.
  • Ignoring security: An agent that can call tools can also cause damage. Tool permissions and sandboxing are mandatory.
  • No monitoring: If you don’t notice when an agent makes wrong decisions, you’ll only find out after harm is done. Audit logging matters.
  • Forgetting human approval: Critical actions need human approval, not blind execution.

Further Reading and Resources on Ollama Agents

Key Takeaways:

  • Ollama is a local model server that supports function calling for AI agents.
  • Most popular frameworks (LangGraph, CrewAI, LangChain) support Ollama.
  • Model choice heavily influences agent quality.
  • MCP connects agents to external tools without requiring you to define each tool yourself.
  • Security, error handling, and human approval are essential for agents with tool access.

FAQ: Ollama for AI Agents - Common Questions

What is Ollama for AI agents?

Ollama is a local model server that serves as a backend for AI agents. It provides a REST API through which agents can request models and use function calling to invoke tools.

Which frameworks support Ollama?

LangGraph, CrewAI, LangChain, LlamaIndex, OpenClaw, and OpenHands support Ollama. Integration typically works through an Ollama class that the framework uses as its model backend.

Does Ollama support function calling?

Yes, since version 0.3.0. You define tools in the API call, the model generates structured tool calls, and your code executes them. Suitable models include Llama 3.1, Qwen 2.5, and Mistral Nemo.

Which model is suitable for AI agents?

For simple agents, Llama 3.1 8B is sufficient. For medium agents with tool selection, Qwen 2.5 14B is better. For complex agents, if your hardware allows, Llama 3.3 70B or Qwen 2.5 32B work well.

Can I use Ollama with MCP?

Ollama itself isn’t an MCP client, but most agent frameworks that support Ollama can also use MCP. LangGraph, CrewAI, and OpenClaw have MCP integrations.

What hardware do I need for Ollama agents?

For an 8B model, a GPU with 8 GB VRAM is enough. For 14B models, you need 12-16 GB VRAM. For 70B models, you need at least 40 GB VRAM, which requires multiple GPUs or a workstation.

How do I secure Ollama agents?

Use tool permissions to restrict what the agent can do, sandboxing to isolate the agent, audit logging to record all actions, and human approval for critical operations.

What does it cost to run Ollama agents?

After your hardware investment, you only pay for electricity. Ollama and most frameworks are open source. Compared to cloud APIs that charge per request, you save significantly if you use agents frequently.

How do I handle errors in agents?

Every tool call needs error handling. When a tool fails, the agent should detect it, retry, or escalate to a human. An agent that crashes on every error isn’t production-ready.

Can I run multiple agents with Ollama?

Yes. With CrewAI or LangGraph, you can build multi-agent systems where multiple agents collaborate. Ollama provides the model for all agents. What matters is ensuring your hardware can handle parallel requests.

References and Further Reading

Back to Blog
Share:

Related Posts