Using the Ollama API
Introduction
Ollama provides a local HTTP API alongside its command-line interface. It lets you access Ollama from your own applications, scripts, and workflows. This is how you integrate locally-run language models into chatbots, agents, and automation systems.
Ollama API - In A Nutshell
After installation, the Ollama server runs in the background and exposes an endpoint at http://localhost:11434. You can query models, manage chat histories, and receive generated responses through it. The API follows a straightforward HTTP pattern that nearly every programming language can handle.
Key Terms and Components
| Term | Definition |
|---|---|
| API endpoint | The address where Ollama listens |
| generate | Endpoint for single text generation requests |
| chat | Endpoint for multi-turn conversations with roles |
| stream | Option to receive responses in real time, chunk by chunk |
| prompt | Text input you send to the model |
When You’ll Use It
The API matters once you move beyond the terminal. That includes custom chat clients, AI agents, automated text processing scripts, or integrations with platforms like n8n.
Starting the Ollama Server
If the server isn’t running, start it with:
ollama serve
The API will then be available at http://localhost:11434. Make sure the model you want to use is already downloaded:
ollama pull llama3
API Request with curl
A simple request without conversation history looks like this:
curl http://localhost:11434/api/generate -d '{
"model": "llama3",
"prompt": "Was ist lokale KI?",
"stream": false
}'
Set "stream": true for streamed responses. Ollama will then return the answer in multiple chunks, which feels more fluid.
API Request with Python
import requests
response = requests.post(
"http://localhost:11434/api/generate",
json={
"model": "llama3",
"prompt": "Erkläre Quantisierung in zwei Sätzen.",
"stream": False
}
)
print(response.json()["response"])
With stream=True, you read the response in a loop:
import requests
import json
response = requests.post(
"http://localhost:11434/api/generate",
json={"model": "llama3", "prompt": "Was ist ein KV-Cache?", "stream": True},
stream=True
)
for line in response.iter_lines():
if line:
data = json.loads(line)
print(data.get("response", ""), end="")
API Request with JavaScript
const response = await fetch('http://localhost:11434/api/generate', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
model: 'llama3',
prompt: 'Was sind Tokens?',
stream: false,
}),
});
const data = await response.json();
console.log(data.response);
Using the Chat Endpoint
The chat endpoint works best when you need conversations with distinct roles:
curl http://localhost:11434/api/chat -d '{
"model": "llama3",
"messages": [
{"role": "user", "content": "Erkläre lokale KI kurz."}
],
"stream": false
}'
This format is compatible with many OpenAI-like clients and is especially useful for building chatbots.
Hardware, Cost, and Security
The API runs locally by default. If you want to expose it across a network or the internet, you’ll need to set up a reverse proxy, configure firewalls, and implement authentication. Learn more in Ollama Network Access.
There are no costs as long as Ollama runs on your own machine. Response speed depends on your hardware. CPU inference is slow; GPU inference is substantially faster.
Key Takeaways
- The Ollama server listens at
http://localhost:11434. - Use
api/generatefor simple prompts andapi/chatfor conversations. - Set
stream: trueto get responses in real time. - The API mimics OpenAI’s design and integrates easily with existing tools.
For setup details, see Installing Ollama. For model management, visit Managing Ollama Models.
FAQ
Do I have to manually start the server every time?
No. On Linux and macOS, Ollama starts automatically as a service after installation. On Windows, it also runs in the background. Only with Docker do you need to launch the container yourself.
Can I expose the API across the network?
Yes, via the environment variable OLLAMA_HOST=0.0.0.0:11434 or through a reverse proxy. Make sure to enable authentication and configure your firewall.
What format do chat messages use?
The chat endpoint expects an array of objects with role and content fields, for example {"role": "user", "content": "Hallo"}.
Can I use Open WebUI with Ollama?
Yes. Open WebUI connects to the local Ollama API by default and provides a modern chat interface in your browser.
How fast is the API?
Speed depends on model size, quantization level, and your hardware. On a GPU with sufficient VRAM, responses often arrive within seconds.
Tools and Resources
curl, Postman, and Insomnia work well for API testing. In Python, use requests or httpx. For JavaScript, a simple fetch call is sufficient.
References
- Ollama API Documentation
- Open WebUI Project Page
- MDN Fetch Documentation


