Skip to content
BotServBotServ
OllamaOpenAIAPIcompatibilityChatGPT

OpenAI-compatible API with Ollama

Use Ollama via /v1 endpoints as OpenAI replacement. Chat, Completions, Embeddings and clients.

S

schutzgeist

3 min read
OpenAI-compatible API with Ollama

OpenAI-Compatible API with Ollama

What this article covers

  • Which endpoints Ollama exposes in OpenAI format.
  • How to migrate existing OpenAI code to Ollama.
  • Key differences and limitations.
  • Examples in Python and curl.
  • Tips for production use.

Introduction: OpenAI-compatible API with Ollama

Ollama provides not just its own API, but also endpoints that match the OpenAI format. This is especially helpful if you’ve already built applications using the OpenAI client library and want to point them at a local Ollama instance instead. Instead of api.openai.com, you simply use the local Ollama address with the /v1 path.

This article walks through which endpoints are available, how the migration works, and what to watch out for.

Key terms

  • OpenAI API: OpenAI’s interface for chat, completions, and embeddings.
  • OpenAI-compatible endpoint: An API that mirrors the OpenAI schema.
  • /v1: The path in Ollama for OpenAI-compatible requests.
  • Chat Completions: Multi-turn conversations.
  • Embeddings: Vector representations of text.
  • API Key: Authentication credential; for Ollama often a placeholder.
  • Base URL: The API address, for example http://localhost:11434/v1.

Available endpoints

Ollama implements many OpenAI endpoints:

  • POST /v1/chat/completions for chat.
  • POST /v1/completions for text completion.
  • POST /v1/embeddings for embeddings.
  • GET /v1/models to list available models.

Not all OpenAI features like functions, tools, or vision are available in every model. Support depends on the model and Ollama version.

Endpoint address

http://localhost:11434/v1

For Ollama running on a different machine:

http://ollama-server.local:11434/v1

Python example with OpenAI client

pip install openai
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama"
)

response = client.chat.completions.create(
    model="llama3.1",
    messages=[
        {"role": "user", "content": "Explain Ollama in three sentences."}
    ]
)

print(response.choices[0].message.content)

Streaming

stream = client.chat.completions.create(
    model="llama3.1",
    messages=[{"role": "user", "content": "Tell me a joke."}],
    stream=True
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Embeddings

response = client.embeddings.create(
    model="nomic-embed-text",
    input="Ollama is a local AI platform."
)

print(response.data[0].embedding[:5])

curl example

curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.1",
    "messages": [
      {"role": "user", "content": "Hello Ollama"}
    ]
  }'

List models

curl http://localhost:11434/v1/models

Differences from the real OpenAI API

  • Authentication: Ollama doesn’t require a real API key, but expects some value.
  • Rate limiting: Usually not present locally.
  • Model names: OpenAI model names must be replaced with Ollama model names.
  • Tools / Functions: Not every model supports Function Calling.
  • Vision: Only certain multimodal models like llava handle images.
  • Streaming: Generally supported.
  • RAG: Must be implemented in your application.

Migrating an existing application

Steps:

  1. Set base_url to http://localhost:11434/v1.
  2. Set api_key to any value like "ollama".
  3. Replace OpenAI model names with Ollama names.
  4. Check which features Ollama doesn’t yet support.
  5. Adjust error handling, as error messages may differ.

When is the OpenAI API useful?

  • You want existing code to run locally with minimal changes.
  • Applications should become privacy-friendly without major refactoring.
  • Comparing OpenAI against local models.
  • Prototypes that will later run on Ollama in production.

When to prefer the native Ollama API?

  • When you need all Ollama features like Modelfiles, Pull, Delete, and Generate.
  • When specific Ollama extensions are required.
  • When the application is being developed originally for Ollama.

Security

  • Ollama runs without authentication by default.
  • Don’t expose it publicly to the internet.
  • Place a reverse proxy with authentication in front if needed.
  • The API key in the client isn’t a real secret, just a placeholder.
  • Run communication over the local network or via VPN/Tailscale.

Common pitfalls

  • Wrong port: Default is 11434.
  • Model not downloaded: Run ollama pull llama3.1 first.
  • Missing api_key: The OpenAI client requires a value, even if Ollama ignores it.
  • Wrong endpoints: /v1/chat/completions, not /api/chat.
  • Vision models: Not every model processes images.
  • Tools unsupported: The model must support Function Calling.
  • CORS issues: Browser requests to localhost can be blocked.

Further reading and resources

FAQ: OpenAI-compatible API with Ollama

Do I need a real OpenAI API key? No, any value like ollama works.

Are all OpenAI features available? No, tools, vision, and certain parameters depend on the model.

Can I use streaming? Yes, streaming via SSE is generally supported.

Which models work? Any local Ollama model accessed by name.

Is the API identical to OpenAI? Almost, but there are limitations and different error messages.

Sources and further reading

Summary: OpenAI-compatible API with Ollama

Ollama’s /v1 endpoints let you run existing OpenAI code locally with minimal changes. By setting base_url to http://localhost:11434/v1 and the API key to a placeholder, you can use models like llama3.1 or qwen2.5 instead of OpenAI models. Keep in mind model availability, missing features like tools or vision, and securing your local Ollama instance. For pure Ollama features, the native API remains the better choice, but for migrations and prototypes, OpenAI compatibility is invaluable.

Back to Blog
Share:

Related Posts