Function Calling with Ollama
What this article covers
- What Function Calling is.
- Which Ollama models support tool use.
- JSON mode and schemas.
- How to structure tool calls in prompts.
- Practical Python example.
Introduction: Function Calling with Ollama
Function Calling allows language models to do more than generate text, they can output structured calls to external functions. This enables a model to fetch data, trigger calculations, or perform actions within an application. Ollama supports Function Calling on compatible models, and with the right prompt structure, a model can return JSON containing function names and parameters.
This article walks through how Function Calling works with Ollama and how to integrate it into your own scripts.
Key concepts
- Function Calling: Model outputs structured function calls.
- Tool Use: Synonym for Function Calling.
- JSON Mode: Model returns exclusively JSON.
- Schema: Definition of allowed functions and their parameters.
- System Prompt: Instructions that prepare the model for Function Calling.
- Named Entity: Recognized entity that becomes a tool argument.
- Tool Result: Output returned to the model after tool execution.
How it works
- Application receives user query.
- Model is provided with available tools.
- Model decides whether to invoke a tool.
- Application executes the tool.
- Result is returned to the model.
- Model formulates response based on the tool result.
Supported models
- Llama 3.1 and newer versions with tool use support.
- Qwen 2.5 and larger variants.
- Mistral, Mixtral, and certain fine-tuned versions.
- Not every model works equally well, as clean JSON output is required.
System prompt for Function Calling
The assistant can call functions when the request requires it.
Respond either with normal text or with a JSON object in the following format when a tool is needed:
{
"function": "tool_name",
"arguments": {
"key": "value"
}
}
Defining available tools
Available tools:
- wetter(location: str) -> str
- rechnen(ausdruck: str) -> str
Python example
import requests
import json
url = "http://localhost:11434/api/generate"
tools = {
"wetter": lambda location: f"Sunny in {location}, 22°C.",
"rechnen": lambda ausdruck: str(eval(ausdruck))
}
system_prompt = """The assistant can call functions.
When a tool is needed, respond exclusively with JSON in this format:
{
"function": "tool_name",
"arguments": {"key": "value"}
}
Available tools:
- wetter(location: str)
- rechnen(ausdruck: str)
"""
def call_ollama(prompt, model="llama3.1"):
response = requests.post(url, json={
"model": model,
"prompt": f"{system_prompt}\n\nUser: {prompt}\nAssistant:",
"format": "json",
"stream": False
})
return response.json()["response"].strip()
def process(user_input):
raw = call_ollama(user_input)
try:
call = json.loads(raw)
name = call["function"]
args = call["arguments"]
result = tools[name](**args)
# Second call with tool result
second = call_ollama(f"Result from {name}({args}): {result}\nNow answer the original question.")
return second
except (json.JSONDecodeError, KeyError, TypeError):
return raw
print(process("What's the weather in Berlin?"))
Using JSON mode
Setting format: json forces Ollama to output JSON. However, the model must be trained to produce valid JSON consistently. With smaller models, specifying the schema in the prompt can help.
Schema in the prompt
{
"function": "tool_name",
"arguments": {
"arg_name": "type"
}
}
Returning tool results
The model needs context to respond to the result:
Result from wetter(location="Berlin"): "Sunny, 22°C."
Answer the question: What's the weather in Berlin?
Error handling
- Invalid JSON: Retry with a shorter prompt.
- Unknown function: Provide feedback to the model.
- Wrong parameter type: Validate before calling the tool.
- No function call: Display the direct text response.
Security
- Never use
evalwithout strict validation. - Validate inputs before passing them to tools.
- Restrict access to sensitive tools.
- Avoid storing secrets in logs.
- Limit permissions of tool functions.
Further reading
FAQ: Function Calling with Ollama
Which models work best? Llama 3.1, Qwen 2.5, and mistral-nemo are solid choices.
Do I need JSON mode? Yes, to ensure reliable and consistent output.
Can Ollama call multiple tools at once? In theory yes, but most models return only one call per response.
What if the JSON is invalid? The tool should return an error and optionally request a retry.
Should I use eval?
Only under extreme restrictions and never with user input.
Sources and further resources
- Ollama API: https://github.com/ollama/ollama/blob/main/docs/api.md
- Function Calling Guide: https://platform.openai.com/docs/guides/function-calling
- llama.cpp Grammars: https://github.com/ggerganov/llama.cpp/blob/master/grammars/README.md
Summary: Function Calling with Ollama
Function Calling extends Ollama with structured tool invocation. With a clear system prompt, tool definitions, JSON mode, and robust error handling, you can automate many workflows. Not every model is suitable, and security matters. When implementing Function Calling, validate inputs, restrict tools, and always have a fallback strategy.


