Evaluating Ollama Models
What this article covers
- Why you should evaluate models
- Which metrics and benchmarks matter
- How to build your own test tasks
- How to compare models automatically
- Tips for fair and meaningful comparisons
Introduction: Evaluating Ollama Models
Ollama offers many models, but not every one performs equally well on every task. One might excel at coding while another shines at mathematics or multilingual summarization. If you evaluate models regularly, you’ll find the right fit for your use case faster and avoid relying on gut feelings.
This article shows how to systematically assess Ollama models, which metrics are worth tracking, and how to design your own test suite.
Key Terms
- Evaluation: Assessment of a model’s performance
- Benchmark: Standardized test for model assessment
- Metric: A measurable property of model behavior
- Perplexity: A measure of model uncertainty
- BLEU / ROUGE: Text generation quality metrics
- Human Evaluation: Manual assessment by people
- Few-Shot: Including a few examples in the prompt
- Zero-Shot: No examples in the prompt
Why Evaluate Models?
- Not every model suits every task
- Larger models aren’t always better
- Quantization affects quality differently across models
- Production applications need reproducible results
- You compare alternatives much faster
Evaluation Dimensions
General Knowledge
Questions from history, geography, science, and general education.
Coding
Programming tasks, code explanations, refactoring, and bug fixes.
Math and Logic
Arithmetic problems, word equations, logical puzzles.
Multilingual Capabilities
Translations, summaries, and responses in different languages.
Factuality and Hallucinations
Testing whether the model invents false information.
Speed and Resources
Tokens per second, RAM and VRAM usage, load times.
Popular Benchmarks
- MMLU: General knowledge across many subjects
- HumanEval: Coding tasks
- GSM8K: Math word problems
- TruthfulQA: Detection of misinformation
- MT-Bench: Multi-turn conversations
You can often run these benchmarks locally using tools like lm-evaluation-harness.
Building Your Own Test Tasks
A simple task collection for your own use case:
[
{
"id": 1,
"category": "coding",
"prompt": "Write a Python function that returns the first n prime numbers.",
"expected": ["def primes(n):", "for", "return"]
},
{
"id": 2,
"category": "summarization",
"prompt": "Summarize the following text in three sentences: ...",
"expected": ["main", "topic"]
}
]
Automated Evaluation in Python
import ollama
import json
def ask(model, prompt):
response = ollama.generate(model=model, prompt=prompt)
return response["response"]
def evaluate(model, tasks):
results = []
for task in tasks:
answer = ask(model, task["prompt"])
score = check_keywords(answer, task["expected"])
results.append({
"id": task["id"],
"category": task["category"],
"score": score
})
return results
def check_keywords(answer, expected):
matches = [kw for kw in expected if kw.lower() in answer.lower()]
return len(matches) / len(expected)
tasks = json.load(open("tasks.json"))
print(evaluate("llama3.1", tasks))
Comparing Models
models = ["llama3.1", "qwen2.5:14b", "mistral"]
for model in models:
scores = evaluate(model, tasks)
avg = sum(r["score"] for r in scores) / len(scores)
print(f"{model}: {avg:.2f}")
Metrics for Generated Text
- Exact Match: Response matches the expected answer exactly
- Keyword Match: Expected terms appear in the response
- ROUGE: Overlap of n-grams between response and reference
- BERTScore: Semantic similarity of answers
- LLM-as-a-Judge: A stronger model evaluates the quality of answers
Measuring Speed
import time
start = time.time()
answer = ask("llama3.1", prompt)
end = time.time()
print(f"Duration: {end - start:.2f} s")
print(f"Tokens per second: {len(answer.split()) / (end - start):.1f}")
For more precise token counts, use ollama ps or check the API response for eval_count and eval_duration.
Tips for Fair Comparisons
- Use identical prompts for all models
- Keep temperature low and consistent, typically around 0.1
- Run on the same hardware and software environment
- Run multiple passes per task
- Draw test cases from your actual use cases
- Document quantization level and timestamp
- Capture speed and resource usage alongside accuracy
Common Pitfalls
- Different prompts: Results become incomparable
- Temperature too high: Answers vary too much
- Too few test tasks: Results lack statistical power
- Relying only on benchmarks: Real-world performance may differ
- Hardware differences: Slower GPU skews timing comparisons
- Not checking hallucinations: Factual errors go undetected
Further Reading and Resources
- BotServ.de Ollama Performance
- BotServ.de Ollama Commands
- BotServ.de Ollama Model Lifecycle
- BotServ.de Local Coding Models
FAQ: Ollama Evaluation
Do I have to evaluate models myself? Not always, but it’s highly recommended for your own use cases.
Which benchmarks matter most? It depends on your use case: MMLU for general knowledge, HumanEval for coding, GSM8K for math.
Can I test two models at the same time? Yes, but not on the same GPU if both are large.
How many test cases do I need? Better 50 than 5, but even a few dozen custom tasks help.
What is LLM-as-a-Judge? Another model evaluates the quality of the answers.
Sources and Further Reading
- EleutherAI lm-evaluation-harness: https://github.com/EleutherAI/lm-evaluation-harness
- HELM: https://crfm.stanford.edu/helm/
- BigBench: https://github.com/google/BIG-bench
- MMLU: https://paperswithcode.com/dataset/mmlu
Summary: Evaluating Ollama Models
Evaluation is your best tool for finding the right Ollama model for a task. Custom test tasks, established benchmarks, and simple metrics like keyword matching or tokens per second provide objective comparisons. What matters most is using consistent conditions, low temperatures, and test coverage that reflects your real-world needs. Regular evaluation lets you make informed choices between model sizes, quantization levels, and architectures.


