Ollama REST API
What This Article Covers
- Available endpoints in Ollama.
- How to generate text, chat, and create embeddings.
- Differences between the native and OpenAI-compatible API.
- Examples using cURL and Python.
- Integration tips and troubleshooting.
Introduction: Ollama REST API
Ollama exposes a powerful REST API that lets external applications and agents interact with locally running models. The API comes in two flavors: the native Ollama API and an OpenAI-compatible endpoint. Both allow you to integrate Ollama into existing tools, scripts, and agent frameworks.
This article walks through the main endpoints, practical examples, and common pitfalls.
Key Terms
- Endpoint: The URL path for an API call.
- Generate: Single-turn text generation from a prompt.
- Chat: Multi-turn conversation with messages.
- Embedding: Converting text into a vector representation.
- Streaming: Responses sent token by token.
- OpenAI-compatible: Endpoint following OpenAI’s schema.
- Model: The identifier of the loaded language model.
API Basics
The API listens on this address by default:
http://localhost:11434
For the OpenAI-compatible endpoint, append /v1:
http://localhost:11434/v1
Listing Available Models
curl http://localhost:11434/api/tags
Response:
{
"models": [
{
"name": "llama3.1:latest",
"model": "llama3.1:latest",
"size": 4928300400,
"parameter_size": "8.0B"
}
]
}
Text Generation
curl http://localhost:11434/api/generate -d '{
"model": "llama3.1",
"prompt": "Explain the Ollama REST API in three sentences.",
"stream": false
}'
Response:
{
"model": "llama3.1",
"response": "The Ollama REST API enables...",
"done": true
}
Chat Endpoint
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the Ollama REST API?"}
],
"stream": false
}'
Response:
{
"model": "llama3.1",
"message": {
"role": "assistant",
"content": "The Ollama REST API is..."
},
"done": true
}
Embeddings
curl http://localhost:11434/api/embed -d '{
"model": "nomic-embed-text",
"input": "This is a sample text."
}'
Response:
{
"model": "nomic-embed-text",
"embeddings": [[0.12, -0.34, 0.56, ...]]
}
OpenAI-Compatible Endpoint
Many tools and SDKs expect the OpenAI schema. Ollama provides:
http://localhost:11434/v1/chat/completions
Example:
curl http://localhost:11434/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "llama3.1",
"messages": [
{"role": "user", "content": "Hello Ollama."}
],
"temperature": 0.7
}'
Python Example
import requests
url = "http://localhost:11434/api/chat"
payload = {
"model": "llama3.1",
"messages": [
{"role": "user", "content": "Explain REST APIs."}
],
"stream": False
}
resp = requests.post(url, json=payload)
print(resp.json()["message"]["content"])
Streaming
To enable true streaming, set stream to true and process incoming lines:
import requests
import json
url = "http://localhost:11434/api/chat"
payload = {
"model": "llama3.1",
"messages": [{"role": "user", "content": "Tell me a joke."}],
"stream": True
}
resp = requests.post(url, json=payload, stream=True)
for line in resp.iter_lines():
if line:
data = json.loads(line)
print(data.get("message", {}).get("content", ""), end="")
Parameters
Common parameters:
- temperature: Controls creativity and randomness.
- top_p: Nucleus sampling parameter.
- top_k: Limits the number of top tokens to sample from.
- num_predict: Maximum number of tokens to generate.
- num_ctx: Context window length.
- repeat_penalty: Penalty for repeated tokens.
- seed: Ensures reproducible outputs.
Network Access
By default, Ollama listens only on 127.0.0.1. To enable network access, set:
export OLLAMA_HOST=0.0.0.0:11434
In Docker:
-e OLLAMA_HOST=0.0.0.0:11434
Note: Ollama does not provide built-in authentication. Secure access via firewall or VPN.
Integration with Other Tools
- Open WebUI:
OLLAMA_BASE_URL=http://localhost:11434 - OpenClaw: native endpoint
http://localhost:11434, not/v1. - LangChain:
OllamaorChatOllamaclasses. - Continue.dev: endpoint
http://localhost:11434/v1with API keyollama.
Common Pitfalls
- Wrong endpoint: OpenClaw requires
/api, not/v1. - Model not loaded: Run
ollama pullfirst. - Insufficient memory: Model fails to load or crashes.
- Not handling streaming:
stream: truereturns multiple JSON lines. - Missing network access: Set
OLLAMA_HOSTand check firewall rules. - OpenAI client refuses connection: Provide API key
ollamaorsk-ollama.
Further Reading and Resources
- BotServ.de Ollama Features
- BotServ.de Ollama Setup
- BotServ.de Open WebUI Administration
- BotServ.de Managing Ollama in Open WebUI
FAQ: Ollama REST API
Which endpoint should I use?
Use /api/chat or /api/generate for Ollama clients. Use /v1/chat/completions for OpenAI-compatible tools.
Do I need an API key?
No, Ollama does not authenticate. Some external tools accept ollama as a dummy key.
Can I use multiple models at once? Yes, but Ollama keeps only one model in memory at a time, the one currently in use.
How do I get embeddings?
Use /api/embed with an embedding model like nomic-embed-text.
Does streaming work?
Yes, set stream: true and process the incoming lines.
References and Further Reading
- Ollama API Docs: https://github.com/ollama/ollama/blob/main/docs/api.md
- OpenAI API Reference: https://platform.openai.com/docs/api-reference
Summary: Ollama REST API
The Ollama REST API provides HTTP access to locally running models. The native API offers generate, chat, and embed endpoints, while the /v1 endpoint provides compatibility with OpenAI clients. Once you understand the endpoint differences, model names, streaming, and parameters, you can seamlessly integrate Ollama into Open WebUI, agents, RAG systems, and custom scripts. The distinction between native and OpenAI-compatible endpoints is especially important to get right.


