Ollama Vision API
What This Article Covers
- How Ollama processes images.
- Which vision models are available.
- How to pass images as Base64 or URL.
- A practical Python example.
- Tips for better image analysis.
Introduction: Ollama Vision API
Ollama supports multimodal models that understand images alongside text. With these models, you can analyze, describe, classify, or summarize images. The API follows OpenAI’s design and lets you send images either as URLs or Base64-encoded data. Vision models are typically larger and slower than text-only models, but they’re invaluable for many tasks.
This article shows how to use the Ollama Vision API.
Key Terms
- Vision model: A model capable of processing images.
- Llava: A popular open-source vision model.
- Base64: Text encoding of binary data.
- Multimodal: Support for multiple input types, such as text and images.
- Embedding: Vector representation of an image.
- Image token: A special token representing image content.
- CLIP: A model that bridges images and text.
- OCR: Optical character recognition in images.
Available Vision Models
- llava
- llava-llama3
- bakllava
- moondream (very small)
- granite3.2-vision
Each model differs in size, speed, and quality.
Python Example with Base64
import requests
import base64
url = "http://localhost:11434/api/chat"
with open("bild.jpg", "rb") as f:
image_base64 = base64.b64encode(f.read()).decode("utf-8")
payload = {
"model": "llava",
"messages": [
{
"role": "user",
"content": "Beschreibe das Bild.",
"images": [image_base64]
}
],
"stream": False
}
response = requests.post(url, json=payload)
print(response.json()["message"]["content"])
Using curl
curl http://localhost:11434/api/chat -d '{
"model": "llava",
"messages": [{
"role": "user",
"content": "Was ist auf dem Bild?",
"images": ["<base64-encoded-image>"]
}],
"stream": false
}'
OpenAI-Compatible Endpoint
import requests
url = "http://localhost:11434/v1/chat/completions"
payload = {
"model": "llava",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Beschreibe das Bild."},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,<base64>"}}
]
}
]
}
response = requests.post(url, json=payload)
print(response.json()["choices"][0]["message"]["content"])
Image URLs
Ollama cannot download images directly from public URLs. Your application must fetch the image and encode it to Base64:
import requests, base64
image_url = "https://example.com/bild.jpg"
img = requests.get(image_url).content
b64 = base64.b64encode(img).decode("utf-8")
Choosing a Model
| Model | Size | Use Case |
|---|---|---|
| moondream | very small | Quick descriptions |
| llava-llama3 | 7B/13B | Good quality |
| bakllava | 7B | Better recognition |
| granite3.2-vision | variable | IBM model |
Tips for Better Image Analysis
- Don’t downscale images too much.
- Ask clear, specific questions.
- Test one model per use case.
- Use high image quality for OCR tasks.
- Multiple images per request are typically not supported.
- Keep context length in mind, as images consume many tokens.
VRAM Considerations
Vision models require more VRAM than text models. A 7B vision model in Q4 quantization can consume 6-8 GB of VRAM. Larger images require additional processing overhead.
Troubleshooting
- Model not responding: Ensure the vision model is loaded.
- Image too large: Resize or compress it.
- Unsupported format: Prefer JPEG or PNG.
- Out of memory: Reduce model size or image resolution.
- Wrong endpoint format: Use the OpenAI-compatible endpoint at
/v1/chat/completions.
Further Reading and Resources
- BotServ.de Ollama Vision
- BotServ.de Ollama Model Recommendations
- BotServ.de Ollama API Troubleshooting
- BotServ.de Local Multimodal AI
FAQ: Ollama Vision API
Can Ollama load images from URLs? No. Your application must download the image and encode it to Base64.
Which vision model is best? Llava-llama3 or Bakllava for quality, Moondream for speed.
How much VRAM does a vision model need? More than a text-only model. A 7B model in Q4 can use 6-8 GB.
Which image formats work? Mostly JPEG and PNG.
Can I send multiple images? The Ollama API typically supports one image per message.
Sources and Further Reading
- Ollama Vision: https://ollama.com/blog/llava
- Llava Paper: https://arxiv.org/abs/2304.08485
- Moondream: https://github.com/vikhyat/moondream
The Ollama Vision API enables image analysis using local multimodal models. Images are passed as Base64 data through either the Ollama chat endpoint or the OpenAI-compatible endpoint. Vision models like Llava, Bakllava, and Moondream offer different levels of quality and speed. By paying attention to image size, VRAM, and model selection, you can reliably analyze images locally.


