Skip to content
BotServBotServ
Tool-CallingAI AgentFunctionsAPIOpenAI FunctionsLocal AI

Tool-Calling

Tool-Calling for AI agents explained: invoke functions, pass parameters, and process responses.

S

schutzgeist

17 min read
Tool-Calling

Tool-Calling

What This Article Covers

  • What tool-calling is and why it’s the difference between a simple chatbot and a true AI agent
  • How the process works, from user input to the final answer
  • How to define your own tool in Python and have the model call it
  • How to provide multiple tools at once and let the model choose
  • How to use tool-calling with local models via Ollama

Introduction: Understanding Tool-Calling

A language model on its own can only generate text. It knows words, patterns, and relationships from its training data, but it cannot query a database, read a file, or send an email. That’s where tool-calling comes in. Tool-calling gives a model the ability to invoke external functions and reach beyond its own knowledge base.

This is the core distinction between a simple chatbot and an AI agent. A chatbot answers questions based on what it has learned. An agent uses tools to fetch current data, perform calculations, or trigger real-world actions. If you want to understand how agents actually work, you need to understand tool-calling. Learn more in What is an AI Agent? and Chatbot vs. AI Agent.

Why Do You Need Tool-Calling?

Imagine asking a language model: “What’s the weather right now in Hamburg?” Without tools, the model has to guess or rely on its training data. But today’s weather wasn’t known when the model was trained. The answer is either wrong or vague.

Tool-calling changes that. You define a function that queries a weather API. The model recognizes it needs this function and returns the function name plus the parameter “Hamburg”. Your code executes the function, fetches the actual weather result, and sends it back to the model. The model then formulates a natural response like “It’s currently 14 degrees in Hamburg with light rain.”

Without tools, a model is confined to its training. With tools, it can access current data, files, databases, and APIs. That’s the critical shift from a text generator to a system that actually does things.

Tool-Calling in a Nutshell

Tool-calling means the model doesn’t directly generate an answer, but instead produces a structured message stating which function it wants to call and what parameters to use. Your program executes the function and sends the result back to the model. The model uses this result to formulate a human-readable response.

The classic example is a weather lookup:

  1. You define a tool get_weather with a parameter location.
  2. The user asks about the weather in Berlin.
  3. The model returns: Call get_weather with location = "Berlin".
  4. Your code executes the function and gets the result, for example “14 degrees, rain”.
  5. You send this result back to the model.
  6. The model formulates the final answer for the user.

The model never executes the function itself. It only tells you which function it needs and what parameters to use. Execution always happens in your code. This is important to understand because it means you maintain full control over what actually happens.

Who Should Use Tool-Calling?

Tool-calling is for developers who want to build more than a simple chatbot. If you’re developing an agent that accesses real data, performs actions, or communicates with other systems, you need tool-calling.

Common scenarios include:

  • Building an agent that fetches current information from the internet.
  • Wanting a model to access your local file storage or database.
  • Developing an assistant that creates calendar entries or sends emails.
  • Using a framework like LangGraph or CrewAI and wanting to understand what happens under the hood. Learn more in Frameworks.

You don’t need years of AI development experience, but you should have basic programming knowledge, ideally in Python. The examples in this article are structured so you can follow them step by step.

Key Concepts in Tool-Calling

TermMeaning
ToolAn external function that the model can call, for example a weather lookup
SchemaThe formal description of a tool including name, parameters, and purpose
Function CallingAnother term for tool-calling, commonly used by OpenAI
ParameterThe arguments the model passes to the tool, for example a location name
Tool CallThe model’s output requesting a function instead of a direct answer
JSONThe format in which tool calls and parameters are structured
APIAn interface through which your tool retrieves data from external services
MCPModel Context Protocol, a standard for binding tools to models uniformly

How Does Tool-Calling Work Step by Step?

Tool-calling follows a clear pattern. Each step matters, and once you understand the principle, you can apply it to any kind of tool.

Step 1: The user asks a question.

The user writes a message, for example “What’s the weather in Munich?” This message goes into the message list that you pass to the model.

Step 2: The model decides if a tool is needed.

The model receives not just the message, but also the list of available tools including their schemas. It analyzes the question and checks whether a tool can help. If yes, it decides to make a tool call. If no, it answers directly.

Step 3: The model returns a tool call.

The tool call is a structured output, usually in JSON format. It contains the function name and parameter values. The model doesn’t generate normal response text, but instead produces this structured instruction.

Step 4: Your code executes the function.

Now it’s your turn. Your program takes the tool call, reads the function name and parameters, and executes the actual function. This could be an API call, a database query, or a calculation.

Step 5: The result goes back to the model.

You take the function result and add it to the message list as a new message, typically with the role “tool”. Then you call the model again.

Step 6: The model formulates the answer.

The model sees the result and formulates a natural response for the user. From “14 degrees, rain” it becomes “It’s currently 14 degrees in Munich with rain.”

This cycle can repeat if the model decides it needs additional tools. For complex tasks, an agent can make multiple tool calls in sequence.

Example: Python with a Simple Tool

Let’s walk through the entire process in Python. We’ll use the OpenAI-compatible API because it’s widely supported by most frameworks and local solutions like Ollama.

import json
from openai import OpenAI

client = OpenAI()

# Step 1: Define the actual function
def get_weather(location: str) -> str:
    # In a real application, you'd call a weather API here.
    # For this example, we return a fixed value.
    return f"Das Wetter in {location} ist sonnig und 22 Grad."

# Step 2: Define the tool schema
tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Gibt das aktuelle Wetter für einen Ort zurück.",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {
                        "type": "string",
                        "description": "Name des Ortes, zum Beispiel 'Berlin'"
                    }
                },
                "required": ["location"]
            }
        }
    }
]

# Step 3: Prepare the user's message
messages = [
    {"role": "user", "content": "Wie ist das Wetter in München?"}
]

# Step 4: Call the model with tools
response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=messages,
    tools=tools,
    tool_choice="auto"
)

# Step 5: Check whether the model wants to call a tool
tool_calls = response.choices[0].message.tool_calls

if tool_calls:
    # Add the model's response to the message list
    messages.append(response.choices[0].message)

    for tool_call in tool_calls:
        # Extract the function name and parameters
        function_name = tool_call.function.name
        arguments = json.loads(tool_call.function.arguments)

        # Execute the function
        if function_name == "get_weather":
            result = get_weather(**arguments)
        else:
            result = "Unbekannte Funktion."

        # Add the result to the message list
        messages.append({
            "role": "tool",
            "tool_call_id": tool_call.id,
            "content": result
        })

    # Step 6: Call the model again so it can formulate the response
    final_response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=messages,
        tools=tools
    )
    print(final_response.choices[0].message.content)
else:
    # If no tool is called, respond directly
    print(response.choices[0].message.content)

Let’s walk through this step by step.

First, we define the get_weather function. It takes a location as a parameter and returns a string. In a real application, you’d call a weather API like OpenWeatherMap here. For this example, we simply return a fixed string.

Next, we define the tool schema. This schema tells the model the function’s name, what it does, and what parameters it expects. The description field is particularly important because the model uses it to decide whether the tool is relevant for a given question. The more precise your description, the better the model’s decision-making.

The messages list contains the user’s question. We pass it along with the tools to the model. The tool_choice parameter set to auto means the model can decide on its own whether to call a tool or respond directly.

After the call, we check whether tool_calls is present in the response. If it is, we append the model’s response to the message list and iterate through each tool call. We extract the function name and parameters, execute the function, and add the result with the role tool to the message list.

Finally, we call the model again, this time with the result in the message list. The model formulates the natural language response and returns it.

Example: Multiple Tools

In practice, you rarely have just one tool. An agent might query weather, determine the current time, and perform calculations. Let’s look at how you define multiple tools and let the model choose between them.

import json
from openai import OpenAI

client = OpenAI()

def get_weather(location: str) -> str:
    return f"Das Wetter in {location} ist bewölkt und 18 Grad."

def get_time(timezone: str) -> str:
    return f"Die aktuelle Zeit in {timezone} ist 14:30 Uhr."

def calculate(expression: str) -> str:
    try:
        result = eval(expression)
        return f"Ergebnis: {result}"
    except Exception:
        return "Ungültiger Ausdruck."

# A dictionary that maps function names to Python functions
available_functions = {
    "get_weather": get_weather,
    "get_time": get_time,
    "calculate": calculate
}

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Gibt das aktuelle Wetter für einen Ort zurück.",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string", "description": "Name des Ortes"}
                },
                "required": ["location"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "get_time",
            "description": "Gibt die aktuelle Uhrzeit für eine Zeitzone zurück.",
            "parameters": {
                "type": "object",
                "properties": {
                    "timezone": {"type": "string", "description": "Zeitzone, zum Beispiel 'Europe/Berlin'"}
                },
                "required": ["timezone"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "calculate",
            "description": "Führt eine mathematische Berechnung durch.",
            "parameters": {
                "type": "object",
                "properties": {
                    "expression": {"type": "string", "description": "Mathematischer Ausdruck, zum Beispiel '2 + 3 * 4'"}
                },
                "required": ["expression"]
            }
        }
    }
]

messages = [
    {"role": "user", "content": "Wie ist das Wetter in Köln und wie spät ist es in Europe/Berlin?"}
]

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=messages,
    tools=tools,
    tool_choice="auto"
)

tool_calls = response.choices[0].message.tool_calls

if tool_calls:
    messages.append(response.choices[0].message)

    for tool_call in tool_calls:
        function_name = tool_call.function.name
        arguments = json.loads(tool_call.function.arguments)

        # Get the function from the dictionary and execute it
        function = available_functions.get(function_name)
        if function:
            result = function(**arguments)
        else:
            result = "Unbekannte Funktion."

        messages.append({
            "role": "tool",
            "tool_call_id": tool_call.id,
            "content": result
        })

    final_response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=messages,
        tools=tools
    )
    print(final_response.choices[0].message.content)
else:
    print(response.choices[0].message.content)

Here we define three functions: get_weather, get_time, and calculate. Each gets its own schema in the tools list. The available_functions dictionary maps function names to actual Python functions. This keeps the code clean because we don’t need a separate if statement for each function.

The user’s question mentions two things: weather and time. The model can decide to request both tools in a single call. The loop over tool_calls then executes both functions and adds both results to the message list. The model formulates a final response that includes both pieces of information.

One important note: each tool’s description must be clear and precise. If two tools have similar descriptions, the model might pick the wrong one. Take time to write good descriptions.

Tool-Calling with Ollama

You don’t need a cloud API to use tool-calling. With Ollama, you can run models locally on your own machine. This is especially relevant when you’re processing sensitive data or want to avoid dependency on external services. Learn more about the motivation at What is local AI?.

Ollama supports tool-calling with specific models, such as llama3.1 or qwen2.5. The API is OpenAI-compatible, which means you can use almost the same code.

import json
import requests

# Ollama runs on localhost:11434 by default
OLLAMA_URL = "http://localhost:11434/api/chat"

def get_weather(location: str) -> str:
    return f"Das Wetter in {location} ist sonnig und 22 Grad."

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Gibt das aktuelle Wetter für einen Ort zurück.",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string", "description": "Name des Ortes"}
                },
                "required": ["location"]
            }
        }
    }
]

messages = [
    {"role": "user", "content": "Wie ist das Wetter in Berlin?"}
]

# Send the initial request to Ollama
response = requests.post(OLLAMA_URL, json={
    "model": "llama3.1",
    "messages": messages,
    "tools": tools,
    "stream": False
})

data = response.json()
message = data.get("message", {})

# Check if the model wants to call a tool
if "tool_calls" in message and message["tool_calls"]:
    messages.append(message)

    for tool_call in message["tool_calls"]:
        function = tool_call["function"]
        function_name = function["name"]
        arguments = function["arguments"]

        if function_name == "get_weather":
            result = get_weather(**arguments)
        else:
            result = "Unbekannte Funktion."

        messages.append({
            "role": "tool",
            "content": result
        })

    # Send the second request with the result
    final_response = requests.post(OLLAMA_URL, json={
        "model": "llama3.1",
        "messages": messages,
        "tools": tools,
        "stream": False
    })

    final_data = final_response.json()
    print(final_data["message"]["content"])
else:
    print(message.get("content", ""))

The structure is almost identical to the OpenAI example. The main difference is that we use requests to call the Ollama API directly instead of the OpenAI Python package. The endpoint URL points to your local Ollama server. The structure of tools and messages remains the same.

Keep in mind that not all local models handle tool-calling equally well. Larger models like llama3.1:8b or qwen2.5:14b are significantly more reliable than very small models. If your model doesn’t format the tool call correctly, try using a larger model or write more precise tool descriptions.

Common Tools for AI Agents with Tool-Calling

As you start building agents, you’ll notice certain tool types appear repeatedly. Here’s an overview of the most common ones with examples.

Web search: The agent searches the internet for current information. Example: A web_search(query) function that calls a search API and returns the top results. Useful for questions about current events or facts not in the model’s training data.

File access: The agent reads or writes files on the local system. Example: read_file(path) reads a file’s contents, write_file(path, content) writes to it. Useful when your agent needs to work with documents on your machine.

Database queries: The agent executes SQL queries. Example: query_database(sql) sends a request to a database and returns results. Useful for agents that need access to structured data.

Email sending: The agent sends emails. Example: send_email(to, subject, body) uses an SMTP server. Useful for assistants that need to dispatch notifications or reports.

Calculator: The agent performs mathematical computations. Example: calculate(expression) evaluates a mathematical expression. Useful because language models are often unreliable at complex arithmetic.

Code execution: The agent runs code. Example: run_python(code) executes a Python script and returns the output. Useful for agents that analyze data or generate scripts.

Each of these tools carries risks. Email sending and code execution are particularly sensitive because they have real-world consequences. Always think carefully about which actions an agent should actually be allowed to perform and what approvals are necessary.

Common Pitfalls with Tool-Calling

Tool-calling seems simple at first, but practice reveals several pitfalls you’ll likely encounter. Here are the most common problems.

The model calls the wrong tool. If tool descriptions are unclear or multiple tools overlap in scope, the model may pick the wrong one. Solution: Write precise, distinct descriptions and avoid overlapping responsibilities.

Parameters are wrong or incomplete. The model might invent parameters not in the schema or omit required ones. Solution: Define required parameters clearly and validate all parameters in your code before executing the function.

The model hallucinates tools. Sometimes the model invents functions that don’t exist. This happens often with small models. Solution: Always check whether the function name exists in your list of available functions before trying to call it.

Small models struggle with tool-calling. Not every model handles tool-calling reliably. Very small models often produce malformed JSON or misunderstand tool descriptions. Solution: Use models explicitly trained for tool-calling and test across different model sizes.

Security risks from uncontrolled execution. If a tool performs sensitive actions like deleting files or sending emails, a wrong tool call can cause real damage. Solution: Limit available tools, validate all parameters, and require human approval for critical actions.

Poor error handling masks problems. If a tool fails and you don’t properly return the error to the model, it can’t respond appropriately. Solution: Catch errors and send a clear error message back to the model so it can decide what to do next.

Hardware, Costs, and Security in Tool-Calling

Hardware: Tool-calling itself demands minimal resources. The computationally intensive part is model inference. If you run a local model through Ollama, you need sufficient RAM and ideally a GPU, depending on model size. An 8B model runs on a modern laptop with 16 GB RAM, while larger models like 14B or 32B require significantly more memory.

Costs: With cloud APIs like OpenAI, you pay per token. Tool-calling means extra API calls because you query the model again after each tool result. With multiple tools per question, costs can add up. With local models via Ollama, there are no API costs, but you bear the hardware costs yourself.

Security: Only enable tools the agent actually needs. A tool that deletes files or executes system-level commands should only run with additional approval. Validate all parameters the model provides before using them. Use sandboxing wherever possible, especially for code execution. Remember that the model never executes code itself, it only suggests what should be executed. Control remains with you.

Further Reading and Resources on Tool Calling

  • What is an AI Agent? - Agent fundamentals and how they differ from chatbots
  • Chatbot vs. AI Agent - The distinction between simple chatbots and true agents
  • Frameworks - Overview of frameworks like LangGraph and CrewAI that abstract tool calling
  • Ollama - Run local models and use tool calling without the cloud
  • What is Local AI? - Why local models matter and what they can do

FAQ - Common Tool Calling Questions

Do all models support tool calling?

No. Models need to be explicitly trained for function calls. Many modern models like Llama 3.1, Qwen 2.5, and Mistral handle it well, but very small or older models often can’t do it reliably. Check the documentation for your specific model.

Can I define multiple tools at once?

Yes. You pass a list of tools, and the model selects the appropriate one or requests multiple calls in a single response. Make sure your tool descriptions are distinct enough that the model doesn’t get confused.

Does a tool have to respond synchronously?

Not necessarily. You can use asynchronous tools as long as the result gets back to the model before the next response is generated. In Python, you can use async functions, but you need to make sure you wait for the result.

How do I secure my tools?

Restrict available tools to what’s necessary, validate all parameters, set permissions, and use sandboxing. Critical actions like sending emails or deleting files should require human approval. Never grant unlimited system access.

What happens if the model hallucinates a tool?

The model invents a function name that doesn’t exist in your tool list. Always check in your code whether the function name exists before trying to call it. If it doesn’t, return an error message to the model so it can try again.

Can a model call multiple tools in sequence?

Yes. For complex tasks, the model can decide after the first tool result that it needs another tool. Implement a loop that continues until the model stops requesting tool calls.

Do I need a framework for tool calling?

No. You can implement tool calling directly with the model provider’s API, as shown in the examples throughout this article. Frameworks like LangGraph or CrewAI save you repetitive work and offer additional features like error handling and loop logic.

Does tool calling work with local models?

Yes, with limitations. Ollama supports tool calling for certain models like Llama 3.1 and Qwen 2.5. Reliability depends on model size. Very small models often produce malformed JSON or call the wrong tools.

How many tools can I define at once?

Theoretically unlimited, but in practice accuracy suffers when you offer too many. The model must process all descriptions and make the right choice. Start with a few tools and expand gradually.

What’s the difference between function calling and tool calling?

The terms are usually used interchangeably. OpenAI traditionally uses “function calling,” while newer documentation and other providers prefer “tool calling.” They mean the same thing: the model indicates which function it wants to call, and your code executes it.

Can the model invent parameters I didn’t define?

Yes, this happens especially with smaller models. Your code should only use parameters defined in the schema and ignore any extras. Validate parameters before passing them to the function.

Sources and Further Reading

  • OpenAI Function Calling Documentation
  • Ollama Tool Support Documentation
  • LangGraph Tools Documentation
  • Model Context Protocol (MCP) Specification
Back to Blog
Share:

Nächster Artikel in AI Agents

Weiterlesen
What Is an AI Agent?

Related Posts