Skip to content
BotServBotServ
OllamaREST APIOpenAIIntegrationLocal AI

Ollama REST API

Call Ollama via REST API. Generate, Chat, Embeddings, Tags and OpenAI-compatible endpoints.

S

schutzgeist

3 min read
Ollama REST API

Ollama REST API

What This Article Covers

  • Available endpoints in Ollama.
  • How to generate text, chat, and create embeddings.
  • Differences between the native and OpenAI-compatible API.
  • Examples using cURL and Python.
  • Integration tips and troubleshooting.

Introduction: Ollama REST API

Ollama exposes a powerful REST API that lets external applications and agents interact with locally running models. The API comes in two flavors: the native Ollama API and an OpenAI-compatible endpoint. Both allow you to integrate Ollama into existing tools, scripts, and agent frameworks.

This article walks through the main endpoints, practical examples, and common pitfalls.

Key Terms

  • Endpoint: The URL path for an API call.
  • Generate: Single-turn text generation from a prompt.
  • Chat: Multi-turn conversation with messages.
  • Embedding: Converting text into a vector representation.
  • Streaming: Responses sent token by token.
  • OpenAI-compatible: Endpoint following OpenAI’s schema.
  • Model: The identifier of the loaded language model.

API Basics

The API listens on this address by default:

http://localhost:11434

For the OpenAI-compatible endpoint, append /v1:

http://localhost:11434/v1

Listing Available Models

curl http://localhost:11434/api/tags

Response:

{
  "models": [
    {
      "name": "llama3.1:latest",
      "model": "llama3.1:latest",
      "size": 4928300400,
      "parameter_size": "8.0B"
    }
  ]
}

Text Generation

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1",
  "prompt": "Explain the Ollama REST API in three sentences.",
  "stream": false
}'

Response:

{
  "model": "llama3.1",
  "response": "The Ollama REST API enables...",
  "done": true
}

Chat Endpoint

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1",
  "messages": [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "What is the Ollama REST API?"}
  ],
  "stream": false
}'

Response:

{
  "model": "llama3.1",
  "message": {
    "role": "assistant",
    "content": "The Ollama REST API is..."
  },
  "done": true
}

Embeddings

curl http://localhost:11434/api/embed -d '{
  "model": "nomic-embed-text",
  "input": "This is a sample text."
}'

Response:

{
  "model": "nomic-embed-text",
  "embeddings": [[0.12, -0.34, 0.56, ...]]
}

OpenAI-Compatible Endpoint

Many tools and SDKs expect the OpenAI schema. Ollama provides:

http://localhost:11434/v1/chat/completions

Example:

curl http://localhost:11434/v1/chat/completions -H "Content-Type: application/json" -d '{
  "model": "llama3.1",
  "messages": [
    {"role": "user", "content": "Hello Ollama."}
  ],
  "temperature": 0.7
}'

Python Example

import requests

url = "http://localhost:11434/api/chat"
payload = {
    "model": "llama3.1",
    "messages": [
        {"role": "user", "content": "Explain REST APIs."}
    ],
    "stream": False
}

resp = requests.post(url, json=payload)
print(resp.json()["message"]["content"])

Streaming

To enable true streaming, set stream to true and process incoming lines:

import requests
import json

url = "http://localhost:11434/api/chat"
payload = {
    "model": "llama3.1",
    "messages": [{"role": "user", "content": "Tell me a joke."}],
    "stream": True
}

resp = requests.post(url, json=payload, stream=True)
for line in resp.iter_lines():
    if line:
        data = json.loads(line)
        print(data.get("message", {}).get("content", ""), end="")

Parameters

Common parameters:

  • temperature: Controls creativity and randomness.
  • top_p: Nucleus sampling parameter.
  • top_k: Limits the number of top tokens to sample from.
  • num_predict: Maximum number of tokens to generate.
  • num_ctx: Context window length.
  • repeat_penalty: Penalty for repeated tokens.
  • seed: Ensures reproducible outputs.

Network Access

By default, Ollama listens only on 127.0.0.1. To enable network access, set:

export OLLAMA_HOST=0.0.0.0:11434

In Docker:

-e OLLAMA_HOST=0.0.0.0:11434

Note: Ollama does not provide built-in authentication. Secure access via firewall or VPN.

Integration with Other Tools

  • Open WebUI: OLLAMA_BASE_URL=http://localhost:11434
  • OpenClaw: native endpoint http://localhost:11434, not /v1.
  • LangChain: Ollama or ChatOllama classes.
  • Continue.dev: endpoint http://localhost:11434/v1 with API key ollama.

Common Pitfalls

  • Wrong endpoint: OpenClaw requires /api, not /v1.
  • Model not loaded: Run ollama pull first.
  • Insufficient memory: Model fails to load or crashes.
  • Not handling streaming: stream: true returns multiple JSON lines.
  • Missing network access: Set OLLAMA_HOST and check firewall rules.
  • OpenAI client refuses connection: Provide API key ollama or sk-ollama.

Further Reading and Resources

FAQ: Ollama REST API

Which endpoint should I use? Use /api/chat or /api/generate for Ollama clients. Use /v1/chat/completions for OpenAI-compatible tools.

Do I need an API key? No, Ollama does not authenticate. Some external tools accept ollama as a dummy key.

Can I use multiple models at once? Yes, but Ollama keeps only one model in memory at a time, the one currently in use.

How do I get embeddings? Use /api/embed with an embedding model like nomic-embed-text.

Does streaming work? Yes, set stream: true and process the incoming lines.

References and Further Reading

Summary: Ollama REST API

The Ollama REST API provides HTTP access to locally running models. The native API offers generate, chat, and embed endpoints, while the /v1 endpoint provides compatibility with OpenAI clients. Once you understand the endpoint differences, model names, streaming, and parameters, you can seamlessly integrate Ollama into Open WebUI, agents, RAG systems, and custom scripts. The distinction between native and OpenAI-compatible endpoints is especially important to get right.

Back to Blog
Share:

Related Posts