Skip to content
BotServBotServ
OllamaVisionAPIImagesLlava

Ollama Vision API

Image analysis with Ollama Vision API. Use Base64, image URLs, Llava and vision models.

S

schutzgeist

3 min read
Ollama Vision API

Ollama Vision API

What This Article Covers

  • How Ollama processes images.
  • Which vision models are available.
  • How to pass images as Base64 or URL.
  • A practical Python example.
  • Tips for better image analysis.

Introduction: Ollama Vision API

Ollama supports multimodal models that understand images alongside text. With these models, you can analyze, describe, classify, or summarize images. The API follows OpenAI’s design and lets you send images either as URLs or Base64-encoded data. Vision models are typically larger and slower than text-only models, but they’re invaluable for many tasks.

This article shows how to use the Ollama Vision API.

Key Terms

  • Vision model: A model capable of processing images.
  • Llava: A popular open-source vision model.
  • Base64: Text encoding of binary data.
  • Multimodal: Support for multiple input types, such as text and images.
  • Embedding: Vector representation of an image.
  • Image token: A special token representing image content.
  • CLIP: A model that bridges images and text.
  • OCR: Optical character recognition in images.

Available Vision Models

  • llava
  • llava-llama3
  • bakllava
  • moondream (very small)
  • granite3.2-vision

Each model differs in size, speed, and quality.

Python Example with Base64

import requests
import base64

url = "http://localhost:11434/api/chat"

with open("bild.jpg", "rb") as f:
    image_base64 = base64.b64encode(f.read()).decode("utf-8")

payload = {
    "model": "llava",
    "messages": [
        {
            "role": "user",
            "content": "Beschreibe das Bild.",
            "images": [image_base64]
        }
    ],
    "stream": False
}

response = requests.post(url, json=payload)
print(response.json()["message"]["content"])

Using curl

curl http://localhost:11434/api/chat -d '{
  "model": "llava",
  "messages": [{
    "role": "user",
    "content": "Was ist auf dem Bild?",
    "images": ["<base64-encoded-image>"]
  }],
  "stream": false
}'

OpenAI-Compatible Endpoint

import requests

url = "http://localhost:11434/v1/chat/completions"

payload = {
    "model": "llava",
    "messages": [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Beschreibe das Bild."},
                {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,<base64>"}}
            ]
        }
    ]
}

response = requests.post(url, json=payload)
print(response.json()["choices"][0]["message"]["content"])

Image URLs

Ollama cannot download images directly from public URLs. Your application must fetch the image and encode it to Base64:

import requests, base64

image_url = "https://example.com/bild.jpg"
img = requests.get(image_url).content
b64 = base64.b64encode(img).decode("utf-8")

Choosing a Model

ModelSizeUse Case
moondreamvery smallQuick descriptions
llava-llama37B/13BGood quality
bakllava7BBetter recognition
granite3.2-visionvariableIBM model

Tips for Better Image Analysis

  • Don’t downscale images too much.
  • Ask clear, specific questions.
  • Test one model per use case.
  • Use high image quality for OCR tasks.
  • Multiple images per request are typically not supported.
  • Keep context length in mind, as images consume many tokens.

VRAM Considerations

Vision models require more VRAM than text models. A 7B vision model in Q4 quantization can consume 6-8 GB of VRAM. Larger images require additional processing overhead.

Troubleshooting

  • Model not responding: Ensure the vision model is loaded.
  • Image too large: Resize or compress it.
  • Unsupported format: Prefer JPEG or PNG.
  • Out of memory: Reduce model size or image resolution.
  • Wrong endpoint format: Use the OpenAI-compatible endpoint at /v1/chat/completions.

Further Reading and Resources

FAQ: Ollama Vision API

Can Ollama load images from URLs? No. Your application must download the image and encode it to Base64.

Which vision model is best? Llava-llama3 or Bakllava for quality, Moondream for speed.

How much VRAM does a vision model need? More than a text-only model. A 7B model in Q4 can use 6-8 GB.

Which image formats work? Mostly JPEG and PNG.

Can I send multiple images? The Ollama API typically supports one image per message.

Sources and Further Reading

The Ollama Vision API enables image analysis using local multimodal models. Images are passed as Base64 data through either the Ollama chat endpoint or the OpenAI-compatible endpoint. Vision models like Llava, Bakllava, and Moondream offer different levels of quality and speed. By paying attention to image size, VRAM, and model selection, you can reliably analyze images locally.

Back to Blog
Share:

Related Posts