Skip to content
BotServBotServ
AI AgentsOllamaLocal AIAgentsAutonomous

Run AI Agents Locally with Ollama

Deploy AI agents locally with Ollama. Setup, tool use, memory, multi-agent systems and practical examples for autonomous agents.

S

schutzgeist

9 min read
Run AI Agents Locally with Ollama

Running AI Agents Locally with Ollama

What This Article Covers

  • How to run AI agents locally using Ollama as your backend.
  • Which agent frameworks work with Ollama and how to set them up.
  • Implementing tool use, memory, and multi-agent systems.
  • Practical examples for autonomous agents: research, code execution, file management.
  • Best practices for security, performance, and reliability.

Introduction: Understanding AI Agents with Ollama

AI agents are programs that leverage a language model to solve tasks independently. They plan, execute tools, observe results, and adjust their approach. With Ollama as a backend, these agents run entirely locally, no cloud, no API costs, no data privacy concerns.

This article is for developers who want to build autonomous agents using local AI. You should be familiar with AI agent basics, how Ollama works, and what Function Calling is. For Python fundamentals, check out IRC-Coding.de.

Why Use Local AI Agents?

Imagine building an agent that sorts your emails, manages appointments, and summarizes documents. In the cloud, every agent step costs money, every API call is logged, every data point flows to a provider. With local AI, everything runs on your hardware, no per-call charges, no data leaving your system.

Local agents shine when you handle sensitive data, need many agent steps (expensive in the cloud), or want independence from cloud providers.

How AI Agents Work with Ollama

An AI agent executes a loop: plan → call tool → observe result → adjust plan. The language model makes decisions; code executes the tools. With Ollama as the backend, the language model runs locally, and tools are local programs (filesystem, shell, web requests).

The core concept: the agent is the loop, Ollama is the brain, tools are the hands.

Who Should Read This?

  • Developers building autonomous agents with local AI.
  • System administrators running agents for automation on their own hardware.
  • Researchers investigating agent behavior using local models.
  • Self-hosters seeking independence from cloud providers.

Prior knowledge of Python, Ollama, and Function Calling is assumed.

Key Concepts for AI Agents with Ollama

  • AI Agent - A program that solves tasks autonomously. Useful for: autonomous workflows.
  • Ollama - Local model server. Useful for: the agent’s brain.
  • Function Calling - Structured AI responses. Useful for: enabling agents to invoke tools.
  • Tool - A function the agent can call. Useful for: the agent’s hands.
  • Memory - The agent’s memory. Useful for: recalling prior steps.
  • MCP - Model Context Protocol. Useful for: a standard for tool integration.
  • OpenHands - Open-source agent platform. Useful for: a ready-made agent with Ollama support.
  • Agent Loop - The core loop: plan → tool → observe → adapt. Useful for: the heart of any agent.

Agent Frameworks with Ollama Support

LangChain with Ollama

LangChain is the most popular framework for LLM applications. It supports Ollama as a backend.

from langchain_community.llms import Ollama
from langchain.agents import AgentExecutor, create_react_agent
from langchain.tools import Tool

# Ollama as backend
llm = Ollama(model="llama3.1", base_url="http://localhost:11434")

# Define tools
def search_wikipedia(query):
    # Implementation here
    return f"Result for: {query}"

tools = [
    Tool(name="Wikipedia", func=search_wikipedia, description="Search Wikipedia")
]

# Create agent
agent = create_react_agent(llm, tools, prompt)
executor = AgentExecutor(agent=agent, tools=tools)

# Run agent
result = executor.invoke({"input": "What is the capital of France?"})
print(result)

LangGraph with Ollama

LangGraph extends LangChain for stateful agents. It allows complex agent graphs with cycles.

from langchain_community.llms import Ollama
from langgraph.graph import StateGraph, END

llm = Ollama(model="llama3.1")

# Define state
from typing import TypedDict, List
class AgentState(TypedDict):
    messages: List[str]
    tool_results: List[str]

# Define nodes
def call_model(state):
    response = llm.invoke(state["messages"][-1])
    return {"messages": state["messages"] + [response]}

def call_tool(state):
    # Tool execution
    return {"tool_results": ["Result"]}

# Build graph
workflow = StateGraph(AgentState)
workflow.add_node("model", call_model)
workflow.add_node("tool", call_tool)
workflow.add_edge("model", "tool")
workflow.add_edge("tool", "model")

CrewAI with Ollama

CrewAI is a framework for multi-agent systems. Multiple agents with different roles collaborate.

from crewai import Agent, Task, Crew
from crewai.llms import Ollama

# Configure LLM
llm = Ollama(model="llama3.1", base_url="http://localhost:11434")

# Define agents
researcher = Agent(
    role="Researcher",
    goal="Gather information",
    backstory="You are an experienced researcher.",
    llm=llm
)

writer = Agent(
    role="Writer",
    goal="Write a report",
    backstory="You are an experienced writer.",
    llm=llm
)

# Define tasks
research_task = Task(description="Research topic X", agent=researcher)
write_task = Task(description="Write a report", agent=writer)

# Create crew
crew = Crew(agents=[researcher, writer], tasks=[research_task, write_task])
result = crew.kickoff()

OpenHands with Ollama

OpenHands (formerly OpenDevin) is a complete agent platform with a web UI. It supports Ollama as a backend.

See Setting up OpenHands for details.

Tool Use with Ollama

Function Calling with Ollama

Ollama supports Function Calling, which agents use to invoke tools.

import requests

def call_ollama_with_tools(messages, tools):
    response = requests.post(
        "http://localhost:11434/api/chat",
        json={
            "model": "llama3.1",
            "messages": messages,
            "tools": tools,
            "stream": False
        }
    )
    return response.json()

# Define tools
tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get weather for a city",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "City name"}
                },
                "required": ["city"]
            }
        }
    }
]

# Agent loop
messages = [{"role": "user", "content": "What is the weather in Berlin?"}]

while True:
    response = call_ollama_with_tools(messages, tools)
    msg = response["message"]

    if msg.get("tool_calls"):
        for tool_call in msg["tool_calls"]:
            # Execute tool
            result = execute_tool(tool_call)
            messages.append(msg)
            messages.append({
                "role": "tool",
                "content": result
            })
    else:
        print(msg["content"])
        break

MCP with Ollama

The Model Context Protocol (MCP) is a standard for tool integration. Ollama models can leverage MCP servers.

See MCP Basics for details.

Memory for Agents

Agents need memory to recall earlier steps. Several memory patterns exist:

Short-Term Memory

The current context, including the most recent messages. Limited by the model’s context window.

messages = [
    {"role": "system", "content": "Du bist ein Agent."},
    {"role": "user", "content": "Schritt 1"},
    {"role": "assistant", "content": "Ergebnis 1"},
    {"role": "user", "content": "Schritt 2"},
    # ...
]

Long-Term Memory

Persistent storage, such as in a vector database. The agent can retrieve earlier insights and findings.

from chromadb import Client

chroma = Client()
collection = chroma.create_collection("agent_memory")

# Store a memory
collection.add(
    documents=["Wichtige Erkenntnis aus Schritt 5"],
    ids=["step_5"]
)

# Retrieve a memory
results = collection.query(query_texts=["Was weiß ich über X?"], n_results=3)

Summary Memory

When context grows too long, the agent summarizes older messages.

def summarize_if_needed(messages, max_length=4000):
    total_length = sum(len(m["content"]) for m in messages)
    if total_length > max_length:
        # Summarize
        summary = call_ollama([
            {"role": "system", "content": "Fasse diese Konversation zusammen."},
            {"role": "user", "content": str(messages)}
        ])
        messages = [
            messages[0],  # Keep system prompt
            {"role": "system", "content": f"Zusammenfassung: {summary}"}
        ]
    return messages

Multi-Agent Setups

Hierarchical Agents

A supervisor agent delegates work to specialized agents:

def supervisor(task):
    # Supervisor decides which agent handles the task
    decision = call_ollama([
        {"role": "system", "content": "Du bist ein Supervisor. Entscheide, welcher Agent dran ist: researcher, coder, oder writer."},
        {"role": "user", "content": task}
    ])

    if "researcher" in decision:
        return researcher_agent(task)
    elif "coder" in decision:
        return coder_agent(task)
    elif "writer" in decision:
        return writer_agent(task)

Parallel Agents

Multiple agents work in parallel on the same problem:

import concurrent.futures

def parallel_agents(task, agents):
    with concurrent.futures.ThreadPoolExecutor() as executor:
        futures = [executor.submit(agent, task) for agent in agents]
        results = [f.result() for f in futures]
    return results

Practical Example 1: Research Agent

An agent that researches a topic, collects sources, and produces a report.

def research_agent(topic):
    messages = [
        {"role": "system", "content": "Du bist ein Recherche-Agent. Nutze die Tools search_web und read_page."},
        {"role": "user", "content": f"Recherchiere: {topic}"}
    ]

    tools = [
        {"type": "function", "function": {"name": "search_web", "parameters": {"query": {"type": "string"}}}},
        {"type": "function", "function": {"name": "read_page", "parameters": {"url": {"type": "string"}}}}
    ]

    while True:
        response = call_ollama_with_tools(messages, tools)
        msg = response["message"]
        messages.append(msg)

        if msg.get("tool_calls"):
            for tc in msg["tool_calls"]:
                result = execute_tool(tc)
                messages.append({"role": "tool", "content": result})
        else:
            return msg["content"]

Practical Example 2: Code Execution Agent

An agent that writes and runs code to solve tasks.

def coder_agent(task):
    messages = [
        {"role": "system", "content": "Du bist ein Code-Agent. Nutze write_code und run_code."},
        {"role": "user", "content": task}
    ]

    tools = [
        {"type": "function", "function": {"name": "write_code", "parameters": {"code": {"type": "string"}}}},
        {"type": "function", "function": {"name": "run_code", "parameters": {"file": {"type": "string"}}}}
    ]

    # Agent loop as above
    # IMPORTANT: Run code in a sandbox!

See Sandbox Environments for security guidance.

Security Considerations

Performance Tips

  • Small Model for Planning: Use a smaller model for planning steps, a larger one for complex reasoning tasks.
  • Watch Context Length: Long agent loops exhaust context windows. Use summary memory.
  • Cache Results: Cache tool outputs to avoid redundant calls.
  • Parallelize: Execute tools in parallel whenever possible.
  • Choose Your Model: For tool use, pick models with strong function calling support, such as llama3.1, qwen2.5, or mistral.

Common Pitfalls

  • No Sandbox: Running code-executing agents without a sandbox is a serious security risk.
  • Oversized Models: Using a 70B model for every agent step wastes VRAM and slows execution.
  • Missing Error Handling: If a tool fails, the agent should adapt, not crash.
  • Exceeding Context Limits: Long loops blow through the context window. Use summary memory.
  • Infinite Loops: Agents can get stuck in cycles. Set a maximum step limit.
  • No Guardrails: Without guardrails, agents can take harmful actions.

Further Reading and Resources on AI Agents with Ollama

Key Takeaways:

  • AI agents with Ollama run entirely local, no cloud required.
  • Frameworks: LangChain, LangGraph, CrewAI, OpenHands.
  • Tool use via function calling or MCP.
  • Memory patterns: short-term, long-term, summary.
  • Multi-agent architectures: hierarchical or parallel.
  • Security essentials: sandboxes, least privilege, guardrails, audit logging.

FAQ: AI Agents with Ollama

What are AI agents?

AI agents are programs that use a language model to solve tasks autonomously. They plan, execute tools, observe results, and adjust their plan accordingly.

Can I use Ollama as a backend for agents?

Yes. Ollama supports Function Calling, which agents need for tool use. Frameworks like LangChain, LangGraph, CrewAI, and OpenHands all support Ollama as a backend.

Which frameworks support Ollama?

LangChain, LangGraph, CrewAI, and OpenHands support Ollama. All of them let you use Ollama as a backend for agents.

How do agents call tools?

Agents use Function Calling to invoke tools. The model returns a structured tool call, your code executes the tool, and sends the result back to the model.

How does memory work for agents?

There are three types: Short-Term Memory (current context), Long-Term Memory (vector database), and Summary Memory (summarized past messages). Which one you pick depends on your use case.

Can I run multiple agents in parallel?

Yes. With frameworks like CrewAI or LangGraph, you can build multi-agent setups where multiple agents with different roles collaborate.

How do I secure AI agents?

Use sandboxed environments for code execution, least privilege for tools, human-in-the-loop for critical actions, guardrails against harmful actions, and audit logging for accountability.

Which model is right for agents?

Models with strong Function Calling: llama3.1, qwen2.5, mistral. For simple tasks, an 8B model works fine. For complex reasoning, use a 32B or 70B model.

What does running local agents cost?

Just hardware costs. Ollama is open source and the models are freely available. You need a machine with enough VRAM for your chosen model.

What should I do about infinite loops?

Set a limit on the number of agent steps. If your agent hasn’t found a solution after N steps, stop it. You can also use Summary Memory to keep the context small.

Resources and Further Reading

Back to Blog
Share:

Related Posts