Skip to content
BotServBotServ
OllamaEvaluationBenchmarkModel ComparisonQuality

Evaluating Ollama Models

Assess local models with benchmarks, metrics, datasets, and comparisons for Ollama.

S

schutzgeist

4 min read
Evaluating Ollama Models

Evaluating Ollama Models

What this article covers

  • Why you should evaluate models
  • Which metrics and benchmarks matter
  • How to build your own test tasks
  • How to compare models automatically
  • Tips for fair and meaningful comparisons

Introduction: Evaluating Ollama Models

Ollama offers many models, but not every one performs equally well on every task. One might excel at coding while another shines at mathematics or multilingual summarization. If you evaluate models regularly, you’ll find the right fit for your use case faster and avoid relying on gut feelings.

This article shows how to systematically assess Ollama models, which metrics are worth tracking, and how to design your own test suite.

Key Terms

  • Evaluation: Assessment of a model’s performance
  • Benchmark: Standardized test for model assessment
  • Metric: A measurable property of model behavior
  • Perplexity: A measure of model uncertainty
  • BLEU / ROUGE: Text generation quality metrics
  • Human Evaluation: Manual assessment by people
  • Few-Shot: Including a few examples in the prompt
  • Zero-Shot: No examples in the prompt

Why Evaluate Models?

  • Not every model suits every task
  • Larger models aren’t always better
  • Quantization affects quality differently across models
  • Production applications need reproducible results
  • You compare alternatives much faster

Evaluation Dimensions

General Knowledge

Questions from history, geography, science, and general education.

Coding

Programming tasks, code explanations, refactoring, and bug fixes.

Math and Logic

Arithmetic problems, word equations, logical puzzles.

Multilingual Capabilities

Translations, summaries, and responses in different languages.

Factuality and Hallucinations

Testing whether the model invents false information.

Speed and Resources

Tokens per second, RAM and VRAM usage, load times.

  • MMLU: General knowledge across many subjects
  • HumanEval: Coding tasks
  • GSM8K: Math word problems
  • TruthfulQA: Detection of misinformation
  • MT-Bench: Multi-turn conversations

You can often run these benchmarks locally using tools like lm-evaluation-harness.

Building Your Own Test Tasks

A simple task collection for your own use case:

[
  {
    "id": 1,
    "category": "coding",
    "prompt": "Write a Python function that returns the first n prime numbers.",
    "expected": ["def primes(n):", "for", "return"]
  },
  {
    "id": 2,
    "category": "summarization",
    "prompt": "Summarize the following text in three sentences: ...",
    "expected": ["main", "topic"]
  }
]

Automated Evaluation in Python

import ollama
import json

def ask(model, prompt):
    response = ollama.generate(model=model, prompt=prompt)
    return response["response"]

def evaluate(model, tasks):
    results = []
    for task in tasks:
        answer = ask(model, task["prompt"])
        score = check_keywords(answer, task["expected"])
        results.append({
            "id": task["id"],
            "category": task["category"],
            "score": score
        })
    return results

def check_keywords(answer, expected):
    matches = [kw for kw in expected if kw.lower() in answer.lower()]
    return len(matches) / len(expected)

tasks = json.load(open("tasks.json"))
print(evaluate("llama3.1", tasks))

Comparing Models

models = ["llama3.1", "qwen2.5:14b", "mistral"]
for model in models:
    scores = evaluate(model, tasks)
    avg = sum(r["score"] for r in scores) / len(scores)
    print(f"{model}: {avg:.2f}")

Metrics for Generated Text

  • Exact Match: Response matches the expected answer exactly
  • Keyword Match: Expected terms appear in the response
  • ROUGE: Overlap of n-grams between response and reference
  • BERTScore: Semantic similarity of answers
  • LLM-as-a-Judge: A stronger model evaluates the quality of answers

Measuring Speed

import time

start = time.time()
answer = ask("llama3.1", prompt)
end = time.time()

print(f"Duration: {end - start:.2f} s")
print(f"Tokens per second: {len(answer.split()) / (end - start):.1f}")

For more precise token counts, use ollama ps or check the API response for eval_count and eval_duration.

Tips for Fair Comparisons

  • Use identical prompts for all models
  • Keep temperature low and consistent, typically around 0.1
  • Run on the same hardware and software environment
  • Run multiple passes per task
  • Draw test cases from your actual use cases
  • Document quantization level and timestamp
  • Capture speed and resource usage alongside accuracy

Common Pitfalls

  • Different prompts: Results become incomparable
  • Temperature too high: Answers vary too much
  • Too few test tasks: Results lack statistical power
  • Relying only on benchmarks: Real-world performance may differ
  • Hardware differences: Slower GPU skews timing comparisons
  • Not checking hallucinations: Factual errors go undetected

Further Reading and Resources

FAQ: Ollama Evaluation

Do I have to evaluate models myself? Not always, but it’s highly recommended for your own use cases.

Which benchmarks matter most? It depends on your use case: MMLU for general knowledge, HumanEval for coding, GSM8K for math.

Can I test two models at the same time? Yes, but not on the same GPU if both are large.

How many test cases do I need? Better 50 than 5, but even a few dozen custom tasks help.

What is LLM-as-a-Judge? Another model evaluates the quality of the answers.

Sources and Further Reading

Summary: Evaluating Ollama Models

Evaluation is your best tool for finding the right Ollama model for a task. Custom test tasks, established benchmarks, and simple metrics like keyword matching or tokens per second provide objective comparisons. What matters most is using consistent conditions, low temperatures, and test coverage that reflects your real-world needs. Regular evaluation lets you make informed choices between model sizes, quantization levels, and architectures.

Back to Blog
Share:

Related Posts