Skip to content
BotServBotServ
OllamaAPIPythonJavaScriptIntegrationSDK

Ollama API Libraries

Use Ollama in Python, JavaScript and more. Official clients, examples, and best practices.

S

schutzgeist

3 min read
Ollama API Libraries

Ollama API Libraries

What This Article Covers

  • Official and community Ollama libraries.
  • Examples in Python and JavaScript.
  • Setting up and using libraries.
  • Differences between native and OpenAI-compatible APIs.
  • Tips for streaming, error handling, and configuration.

Introduction: Ollama API Libraries

Ollama exposes a REST API accessible from virtually any programming language. Python and JavaScript benefit from official or popular client libraries that streamline requests, streaming, and embeddings. If you want to integrate Ollama into your own applications, agents, or scripts, the right library can save considerable time.

This article walks through available libraries, how to install them, and common patterns you’ll use with each.

Key Terms

  • Client: Library that abstracts API calls.
  • SDK: Software Development Kit; includes libraries and tools.
  • REST API: HTTP-based interface.
  • OpenAI-compatible API: Endpoint that mirrors the OpenAI schema.
  • Streaming: Responses arrive incrementally, chunk by chunk.
  • Chat Completion: Multi-turn conversation.
  • Embedding: Vector representation of text.

Official Python Library

Install the ollama library via pip:

pip install ollama

Single Response

import ollama

response = ollama.chat(
    model="llama3.1",
    messages=[
        {"role": "user", "content": "Explain Ollama APIs."}
    ]
)
print(response["message"]["content"])

Streaming

import ollama

stream = ollama.chat(
    model="llama3.1",
    messages=[{"role": "user", "content": "Tell me a joke."}],
    stream=True
)

for chunk in stream:
    print(chunk["message"]["content"], end="")

Embeddings

import ollama

response = ollama.embed(
    model="nomic-embed-text",
    input="Ollama is a tool for local AI."
)
print(response["embeddings"][0][:5])

Generate

import ollama

response = ollama.generate(
    model="llama3.1",
    prompt="Explain quantization in three sentences."
)
print(response["response"])

JavaScript / TypeScript

For Node.js, use the ollama package:

npm install ollama

Example

import ollama from 'ollama';

const response = await ollama.chat({
  model: 'llama3.1',
  messages: [{ role: 'user', content: 'Explain Ollama APIs.' }],
});

console.log(response.message.content);

Streaming

import ollama from 'ollama';

const stream = await ollama.chat({
  model: 'llama3.1',
  messages: [{ role: 'user', content: 'Tell me a joke.' }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.message.content);
}

Other Languages

Many languages simply use HTTP requests. Here are a few examples:

curl

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1",
  "messages": [{"role": "user", "content": "Hello"}]
}'

Python with requests

import requests

resp = requests.post(
    "http://localhost:11434/api/chat",
    json={
        "model": "llama3.1",
        "messages": [{"role": "user", "content": "Hello"}],
        "stream": False
    }
)
print(resp.json()["message"]["content"])

OpenAI-Compatible Client

Many libraries accept a base_url parameter to use Ollama instead of OpenAI:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama"
)

resp = client.chat.completions.create(
    model="llama3.1",
    messages=[{"role": "user", "content": "Hello"}]
)
print(resp.choices[0].message.content)

Configuration

Most clients allow you to customize the host and port:

import ollama

client = ollama.Client(host="http://ollama-server.local:11434")

Or use an environment variable:

export OLLAMA_HOST="http://ollama-server.local:11434"

Error Handling

Common issues you might encounter:

  • Connection failed: Ollama is not running or the host is incorrect.
  • Model not found: Model name is wrong or hasn’t been downloaded.
  • Timeout: Response is taking too long; check GPU utilization.
  • Streaming breaks: Connection issue or chunk processing problem.

Simple error handling example:

import ollama

try:
    response = ollama.generate(model="llama3.1", prompt="Hello")
    print(response["response"])
except ollama.ResponseError as e:
    print("Error:", e.error)

Tips

  • Use the native ollama library when you need all Ollama features.
  • Use OpenAI-compatible clients if you already have OpenAI code.
  • Enable streaming for better UX in interactive applications.
  • Set API keys if Ollama sits behind a proxy that requires authentication.
  • Centralize connection parameters.

Further Reading and Resources

FAQ: Ollama API Libraries

Is there an official Python library? Yes, install it with pip install ollama.

Can I use OpenAI clients with Ollama? Yes, through the OpenAI-compatible API at /v1.

Do I need an API key? No, Ollama itself doesn’t authenticate. Proxies may require one.

How do I use streaming? Set the stream=True parameter and iterate over chunks.

Which languages are supported? Any language that can send HTTP requests. Official clients exist for Python and JavaScript.

Sources and Further Reading

Summary: Ollama API Libraries

Ollama integrates into applications via official Python and JavaScript libraries or simple HTTP requests. The native API exposes all Ollama features, while the OpenAI-compatible API simplifies migration of existing projects. Correct host configuration, accurate model names, and proper error handling are essential. When you leverage streaming, embeddings, and chat completion correctly, you build responsive, privacy-respecting AI applications on Ollama.

Back to Blog
Share:

Related Posts