OpenAI-Compatible API with Ollama
What this article covers
- Which endpoints Ollama exposes in OpenAI format.
- How to migrate existing OpenAI code to Ollama.
- Key differences and limitations.
- Examples in Python and curl.
- Tips for production use.
Introduction: OpenAI-compatible API with Ollama
Ollama provides not just its own API, but also endpoints that match the OpenAI format. This is especially helpful if you’ve already built applications using the OpenAI client library and want to point them at a local Ollama instance instead. Instead of api.openai.com, you simply use the local Ollama address with the /v1 path.
This article walks through which endpoints are available, how the migration works, and what to watch out for.
Key terms
- OpenAI API: OpenAI’s interface for chat, completions, and embeddings.
- OpenAI-compatible endpoint: An API that mirrors the OpenAI schema.
/v1: The path in Ollama for OpenAI-compatible requests.- Chat Completions: Multi-turn conversations.
- Embeddings: Vector representations of text.
- API Key: Authentication credential; for Ollama often a placeholder.
- Base URL: The API address, for example
http://localhost:11434/v1.
Available endpoints
Ollama implements many OpenAI endpoints:
POST /v1/chat/completionsfor chat.POST /v1/completionsfor text completion.POST /v1/embeddingsfor embeddings.GET /v1/modelsto list available models.
Not all OpenAI features like functions, tools, or vision are available in every model. Support depends on the model and Ollama version.
Endpoint address
http://localhost:11434/v1
For Ollama running on a different machine:
http://ollama-server.local:11434/v1
Python example with OpenAI client
pip install openai
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama"
)
response = client.chat.completions.create(
model="llama3.1",
messages=[
{"role": "user", "content": "Explain Ollama in three sentences."}
]
)
print(response.choices[0].message.content)
Streaming
stream = client.chat.completions.create(
model="llama3.1",
messages=[{"role": "user", "content": "Tell me a joke."}],
stream=True
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Embeddings
response = client.embeddings.create(
model="nomic-embed-text",
input="Ollama is a local AI platform."
)
print(response.data[0].embedding[:5])
curl example
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.1",
"messages": [
{"role": "user", "content": "Hello Ollama"}
]
}'
List models
curl http://localhost:11434/v1/models
Differences from the real OpenAI API
- Authentication: Ollama doesn’t require a real API key, but expects some value.
- Rate limiting: Usually not present locally.
- Model names: OpenAI model names must be replaced with Ollama model names.
- Tools / Functions: Not every model supports Function Calling.
- Vision: Only certain multimodal models like
llavahandle images. - Streaming: Generally supported.
- RAG: Must be implemented in your application.
Migrating an existing application
Steps:
- Set
base_urltohttp://localhost:11434/v1. - Set
api_keyto any value like"ollama". - Replace OpenAI model names with Ollama names.
- Check which features Ollama doesn’t yet support.
- Adjust error handling, as error messages may differ.
When is the OpenAI API useful?
- You want existing code to run locally with minimal changes.
- Applications should become privacy-friendly without major refactoring.
- Comparing OpenAI against local models.
- Prototypes that will later run on Ollama in production.
When to prefer the native Ollama API?
- When you need all Ollama features like Modelfiles, Pull, Delete, and Generate.
- When specific Ollama extensions are required.
- When the application is being developed originally for Ollama.
Security
- Ollama runs without authentication by default.
- Don’t expose it publicly to the internet.
- Place a reverse proxy with authentication in front if needed.
- The API key in the client isn’t a real secret, just a placeholder.
- Run communication over the local network or via VPN/Tailscale.
Common pitfalls
- Wrong port: Default is
11434. - Model not downloaded: Run
ollama pull llama3.1first. - Missing
api_key: The OpenAI client requires a value, even if Ollama ignores it. - Wrong endpoints:
/v1/chat/completions, not/api/chat. - Vision models: Not every model processes images.
- Tools unsupported: The model must support Function Calling.
- CORS issues: Browser requests to
localhostcan be blocked.
Further reading and resources
- BotServ.de Ollama REST API
- BotServ.de Ollama API libraries
- BotServ.de Ollama Security
- BotServ.de Docker Reverse Proxy
FAQ: OpenAI-compatible API with Ollama
Do I need a real OpenAI API key?
No, any value like ollama works.
Are all OpenAI features available? No, tools, vision, and certain parameters depend on the model.
Can I use streaming? Yes, streaming via SSE is generally supported.
Which models work? Any local Ollama model accessed by name.
Is the API identical to OpenAI? Almost, but there are limitations and different error messages.
Sources and further reading
- Ollama OpenAI Compatibility: https://ollama.com/blog/openai-compatibility
- OpenAI API Reference: https://platform.openai.com/docs/api-reference
- Ollama API Docs: https://github.com/ollama/ollama/blob/main/docs/api.md
Summary: OpenAI-compatible API with Ollama
Ollama’s /v1 endpoints let you run existing OpenAI code locally with minimal changes. By setting base_url to http://localhost:11434/v1 and the API key to a placeholder, you can use models like llama3.1 or qwen2.5 instead of OpenAI models. Keep in mind model availability, missing features like tools or vision, and securing your local Ollama instance. For pure Ollama features, the native API remains the better choice, but for migrations and prototypes, OpenAI compatibility is invaluable.


