Benchmarking Ollama Models
What This Article Covers
- Why benchmarking matters.
- Which metrics to measure.
- How to calculate tokens per second.
- How to measure VRAM and RAM.
- Comparing models with different quantization levels.
Introduction: Benchmarking Ollama Models
Local models come in many sizes, quantization levels, and architectures. Not every model runs at the same speed on every piece of hardware, nor does every model produce equivalent quality. Benchmarks help you determine which model fits your specific use case and hardware. Three factors matter most: speed, memory consumption, and response quality.
This article shows how to benchmark Ollama models yourself without hassle.
Key Terms
- Throughput: Number of tokens per second.
- Latency: Time until the first response arrives.
- TTFT: Time To First Token.
- VRAM: GPU video memory.
- Quantization: Reducing model precision.
- Prompt tokens: Tokens in the input text.
- Completion tokens: Tokens in the response text.
- Perplexity: A measure of model uncertainty.
What to Measure
- Tokens per second (total): Overall speed.
- Prompt eval time: Duration to process context.
- Eval time: Duration to generate response tokens.
- VRAM usage: Is enough VRAM available?
- RAM usage: For CPU-only operation.
- Quality: Test responses against prepared questions.
Simple Python Benchmark
import requests
import time
url = "http://localhost:11434/api/generate"
prompt = "Erkläre kurz, was ein neuronales Netz ist."
model = "llama3.1"
payload = {
"model": model,
"prompt": prompt,
"stream": False,
"options": {
"num_predict": 100
}
}
start = time.time()
response = requests.post(url, json=payload)
end = time.time()
data = response.json()
duration = end - start
tokens = data.get("eval_count", 0)
tps = tokens / duration if duration > 0 else 0
print(f"Dauer: {duration:.2f} s")
print(f"Tokens: {tokens}")
print(f"Tokens/s: {tps:.2f}")
print(f"VRAM: prüfe mit nvidia-smi")
Examining the Response
print(data["response"])
Comparing Multiple Models
models = ["llama3.1", "qwen2.5:7b", "mistral:7b"]
prompt = "Was ist Docker?"
for model in models:
payload = {
"model": model,
"prompt": prompt,
"stream": False,
"options": {"num_predict": 100}
}
start = time.time()
r = requests.post(url, json=payload)
end = time.time()
data = r.json()
duration = end - start
tokens = data.get("eval_count", 0)
print(f"{model}: {tokens/duration:.2f} tokens/s")
Prompt Eval vs. Completion Eval
When using Ollama with stream: false, the response includes fields like prompt_eval_count and eval_count. You can derive more granular metrics from these:
- prompt eval time: Duration to process the context.
- eval time: Duration to generate the response.
Measuring VRAM
nvidia-smi --query-gpu=memory.used --format=csv -l 1
For AMD:
rocm-smi --showmeminfo vram
Test Questions for Quality
Speed benchmarks alone tell you nothing about quality. Prepared questions help evaluate actual performance:
- Summarizing a text passage.
- Explaining a concept.
- Translation tasks.
- Code generation.
- Logical puzzles.
Comparing Quantization Levels
| Quant | Tokens/s | Quality | VRAM |
|---|---|---|---|
| Q4_0 | high | lower | low |
| Q4_K_M | high | good | low |
| Q5_K_M | moderate | very good | moderate |
| Q8_0 | lower | high | high |
Recommendation: Find the smallest quantization that still meets your quality requirements.
Automated Benchmark
import json
import time
import requests
def benchmark(model, prompt, num_predict=100):
url = "http://localhost:11434/api/generate"
start = time.time()
r = requests.post(url, json={
"model": model,
"prompt": prompt,
"stream": False,
"options": {"num_predict": num_predict}
})
end = time.time()
data = r.json()
return {
"model": model,
"duration": end - start,
"tokens": data.get("eval_count", 0),
"tps": data.get("eval_count", 0) / (end - start)
}
results = [benchmark(m, "Erkläre KI in drei Sätzen.") for m in ["llama3.1", "qwen2.5"]]
print(json.dumps(results, indent=2))
Tips
- Restart Ollama before each run to prevent cache interference from prior operations.
- Use the same prompt and
num_predictvalue across all tests. - Run multiple iterations and calculate the average.
- Set temperature to 0 for reproducible results.
- Monitor hardware temperature to avoid throttling.
Common Pitfalls
- First run is slower: The model needs to load.
- Varying prompt length: Affects timing measurements.
- Streaming: Synchronous measurement becomes difficult.
- GPU throttling: Performance degrades after sustained load.
- Wrong quantization: Model doesn’t fit in VRAM.
- Cache effects: Model persists in memory between runs.
Further Reading
- BotServ.de Ollama Performance
- BotServ.de Ollama Quantization in Practice
- BotServ.de Ollama Model Updates
- BotServ.de Ollama Evaluation
FAQ: Ollama Benchmarks
What’s a good tokens-per-second figure? It depends on the model and hardware. For local 7B models, 10-30 tokens/s is typical.
Should I always use the highest quantization? No. If quality is acceptable for your use case, a lower quantization is faster.
How do I measure VRAM?
Use nvidia-smi or rocm-smi while the model is running.
Why is the second run faster? Because the model is already loaded in memory.
Can I benchmark quality automatically? Only to a limited degree. Prepared questions and manual review work best.
Sources and Further Reading
- Ollama API: https://github.com/ollama/ollama/blob/main/docs/api.md
- llama.cpp Benchmarks: https://github.com/ggerganov/llama.cpp/blob/master/examples/perplexity/README.md
- Token Counting: https://platform.openai.com/tokenizer
Benchmarking is essential for finding the right Ollama models for your hardware and use cases. Measure tokens per second, TTFT, VRAM and RAM consumption, and response quality. Simple Python scripts let you systematically compare models and quantization levels. By running multiple iterations with identical prompts, you’ll get reproducible results and can deliberately choose the best balance between speed and quality.


