Ollama Model Comparison
What this article covers
- Key criteria for comparing models.
- How to evaluate speed and quality dimensions.
- Benchmarking methods.
- Practical comparison workflows.
- Tips for meaningful results.
Introduction: Ollama Model Comparison
The Ollama library grows constantly. Models differ in size, speed, capabilities, and hardware requirements. Anyone looking for the right model for their use case should move beyond published benchmark numbers and run their own tests. A systematic comparison helps you find the best model for your hardware and specific application.
This article walks through how to compare Ollama models effectively.
Key terms
- Benchmark: Standardized test.
- Latency: Time until the first response arrives.
- Throughput: Tokens per second.
- Accuracy: Correctness of answers.
- Hallucination: Fabricated information.
- Quant: Quantization variant.
- Task-Specific: Specialized task.
- Gold-Standard: Reference answer.
Comparison criteria
| Criterion | What it means |
|---|---|
| Speed | Tokens per second, startup time |
| Quality | Correctness, completeness |
| VRAM requirement | Memory footprint |
| Context | Maximum token length |
| Language | German, English, multilingual |
| Specialization | Coding, chat, vision, reasoning |
| License | Commercial use allowed |
Workflow
- Define your use case.
- Select candidate models.
- Create test tasks.
- Evaluate each candidate fairly.
- Measure speed.
- Assess quality.
- Document VRAM and RAM needs.
- Choose the best trade-off.
Measuring speed
ollama run llama3.1:8b --verbose
Or with a script:
import time, requests
for model in ["llama3.1:8b", "qwen2.5:7b"]:
start = time.time()
response = requests.post("http://localhost:11434/api/generate", json={
"model": model,
"prompt": "Erkläre Docker.",
"stream": False
})
duration = time.time() - start
eval_count = response.json().get("eval_count", 0)
print(f"{model}: {eval_count/duration:.2f} tokens/s")
Testing quality
Use prepared questions with known answers:
- Summarization.
- Translation.
- Math.
- Factual retrieval.
- Coding task.
- Logic puzzle.
Check each answer against a gold-standard or manual evaluation.
Keep prompts identical
ollama run llama3.1:8b "Summarize Docker in three sentences."
ollama run qwen2.5:7b "Summarize Docker in three sentences."
Parameter normalization
Use the same temperature, num_predict, and num_ctx for all candidates:
{
"options": {
"temperature": 0.1,
"num_predict": 200,
"num_ctx": 4096
}
}
Tips
- Run multiple rounds per model.
- Keep prompts identical.
- Evaluate models for your specific use case.
- Watch for hallucinations.
- Note the quantization for each model.
- Don’t overweight a single response.
- Avoid bias.
Tools
ollamacommands.curlandjq.- Python scripts.
- VRAM calculators.
- MTEB for embeddings.
- LMSYS Chatbot Arena as reference.
Further reading
- BotServ.de Ollama Model Benchmarks
- BotServ.de Ollama Model Recommendations
- BotServ.de Ollama Model Gallery
- BotServ.de Tools Model Finder
FAQ: Ollama Model Comparison
How many models should I test? Two to five good candidates are usually enough.
Which model is objectively the best? There isn’t one. It depends on your application and hardware.
Do I need benchmarks? Yes, but your own tasks often give more relevant results.
How do I compare speed? Measure tokens per second using the same prompt.
Should I test different quantizations? Yes. Q4 versus Q5 can make significant differences in quality and speed.
Sources and further reading
- LMSYS Arena: https://chat.lmsys.org/
- MTEB: https://huggingface.co/spaces/mteb/leaderboard
- Ollama Library: https://ollama.com/library
Summary: Ollama Model Comparison
A good model comparison considers speed, quality, VRAM footprint, context length, and specialization. Identical prompts and parameters are essential. Benchmarks help, but testing on your own tasks yields the most realistic results. Testing two to five strong candidates on your own workloads helps you find the best trade-off for your hardware and project.


