Skip to content
BotServBotServ
OllamaModelsComparisonBenchmarkTesting

Ollama Model Comparison

Compare Ollama models systematically. Criteria, benchmarks, tasks, and practical testing methods.

S

schutzgeist

2 min read
Ollama Model Comparison

Ollama Model Comparison

What this article covers

  • Key criteria for comparing models.
  • How to evaluate speed and quality dimensions.
  • Benchmarking methods.
  • Practical comparison workflows.
  • Tips for meaningful results.

Introduction: Ollama Model Comparison

The Ollama library grows constantly. Models differ in size, speed, capabilities, and hardware requirements. Anyone looking for the right model for their use case should move beyond published benchmark numbers and run their own tests. A systematic comparison helps you find the best model for your hardware and specific application.

This article walks through how to compare Ollama models effectively.

Key terms

  • Benchmark: Standardized test.
  • Latency: Time until the first response arrives.
  • Throughput: Tokens per second.
  • Accuracy: Correctness of answers.
  • Hallucination: Fabricated information.
  • Quant: Quantization variant.
  • Task-Specific: Specialized task.
  • Gold-Standard: Reference answer.

Comparison criteria

CriterionWhat it means
SpeedTokens per second, startup time
QualityCorrectness, completeness
VRAM requirementMemory footprint
ContextMaximum token length
LanguageGerman, English, multilingual
SpecializationCoding, chat, vision, reasoning
LicenseCommercial use allowed

Workflow

  1. Define your use case.
  2. Select candidate models.
  3. Create test tasks.
  4. Evaluate each candidate fairly.
  5. Measure speed.
  6. Assess quality.
  7. Document VRAM and RAM needs.
  8. Choose the best trade-off.

Measuring speed

ollama run llama3.1:8b --verbose

Or with a script:

import time, requests

for model in ["llama3.1:8b", "qwen2.5:7b"]:
    start = time.time()
    response = requests.post("http://localhost:11434/api/generate", json={
        "model": model,
        "prompt": "Erkläre Docker.",
        "stream": False
    })
    duration = time.time() - start
    eval_count = response.json().get("eval_count", 0)
    print(f"{model}: {eval_count/duration:.2f} tokens/s")

Testing quality

Use prepared questions with known answers:

  • Summarization.
  • Translation.
  • Math.
  • Factual retrieval.
  • Coding task.
  • Logic puzzle.

Check each answer against a gold-standard or manual evaluation.

Keep prompts identical

ollama run llama3.1:8b "Summarize Docker in three sentences."
ollama run qwen2.5:7b "Summarize Docker in three sentences."

Parameter normalization

Use the same temperature, num_predict, and num_ctx for all candidates:

{
  "options": {
    "temperature": 0.1,
    "num_predict": 200,
    "num_ctx": 4096
  }
}

Tips

  • Run multiple rounds per model.
  • Keep prompts identical.
  • Evaluate models for your specific use case.
  • Watch for hallucinations.
  • Note the quantization for each model.
  • Don’t overweight a single response.
  • Avoid bias.

Tools

  • ollama commands.
  • curl and jq.
  • Python scripts.
  • VRAM calculators.
  • MTEB for embeddings.
  • LMSYS Chatbot Arena as reference.

Further reading

FAQ: Ollama Model Comparison

How many models should I test? Two to five good candidates are usually enough.

Which model is objectively the best? There isn’t one. It depends on your application and hardware.

Do I need benchmarks? Yes, but your own tasks often give more relevant results.

How do I compare speed? Measure tokens per second using the same prompt.

Should I test different quantizations? Yes. Q4 versus Q5 can make significant differences in quality and speed.

Sources and further reading

Summary: Ollama Model Comparison

A good model comparison considers speed, quality, VRAM footprint, context length, and specialization. Identical prompts and parameters are essential. Benchmarks help, but testing on your own tasks yields the most realistic results. Testing two to five strong candidates on your own workloads helps you find the best trade-off for your hardware and project.

Back to Blog
Share:

Related Posts