Skip to content
BotServBotServ
OllamaBenchmarkPerformanceTokensMetrics

Benchmark Ollama Models

Benchmark Ollama models with custom scripts and tools. Check tokens per second, VRAM, and quality.

S

schutzgeist

3 min read
Benchmark Ollama Models

Benchmarking Ollama Models

What This Article Covers

  • Why benchmarking matters.
  • Which metrics to measure.
  • How to calculate tokens per second.
  • How to measure VRAM and RAM.
  • Comparing models with different quantization levels.

Introduction: Benchmarking Ollama Models

Local models come in many sizes, quantization levels, and architectures. Not every model runs at the same speed on every piece of hardware, nor does every model produce equivalent quality. Benchmarks help you determine which model fits your specific use case and hardware. Three factors matter most: speed, memory consumption, and response quality.

This article shows how to benchmark Ollama models yourself without hassle.

Key Terms

  • Throughput: Number of tokens per second.
  • Latency: Time until the first response arrives.
  • TTFT: Time To First Token.
  • VRAM: GPU video memory.
  • Quantization: Reducing model precision.
  • Prompt tokens: Tokens in the input text.
  • Completion tokens: Tokens in the response text.
  • Perplexity: A measure of model uncertainty.

What to Measure

  • Tokens per second (total): Overall speed.
  • Prompt eval time: Duration to process context.
  • Eval time: Duration to generate response tokens.
  • VRAM usage: Is enough VRAM available?
  • RAM usage: For CPU-only operation.
  • Quality: Test responses against prepared questions.

Simple Python Benchmark

import requests
import time

url = "http://localhost:11434/api/generate"

prompt = "Erkläre kurz, was ein neuronales Netz ist."
model = "llama3.1"

payload = {
    "model": model,
    "prompt": prompt,
    "stream": False,
    "options": {
        "num_predict": 100
    }
}

start = time.time()
response = requests.post(url, json=payload)
end = time.time()

data = response.json()
duration = end - start
tokens = data.get("eval_count", 0)
tps = tokens / duration if duration > 0 else 0

print(f"Dauer: {duration:.2f} s")
print(f"Tokens: {tokens}")
print(f"Tokens/s: {tps:.2f}")
print(f"VRAM: prüfe mit nvidia-smi")

Examining the Response

print(data["response"])

Comparing Multiple Models

models = ["llama3.1", "qwen2.5:7b", "mistral:7b"]
prompt = "Was ist Docker?"

for model in models:
    payload = {
        "model": model,
        "prompt": prompt,
        "stream": False,
        "options": {"num_predict": 100}
    }
    start = time.time()
    r = requests.post(url, json=payload)
    end = time.time()
    data = r.json()
    duration = end - start
    tokens = data.get("eval_count", 0)
    print(f"{model}: {tokens/duration:.2f} tokens/s")

Prompt Eval vs. Completion Eval

When using Ollama with stream: false, the response includes fields like prompt_eval_count and eval_count. You can derive more granular metrics from these:

  • prompt eval time: Duration to process the context.
  • eval time: Duration to generate the response.

Measuring VRAM

nvidia-smi --query-gpu=memory.used --format=csv -l 1

For AMD:

rocm-smi --showmeminfo vram

Test Questions for Quality

Speed benchmarks alone tell you nothing about quality. Prepared questions help evaluate actual performance:

  • Summarizing a text passage.
  • Explaining a concept.
  • Translation tasks.
  • Code generation.
  • Logical puzzles.

Comparing Quantization Levels

QuantTokens/sQualityVRAM
Q4_0highlowerlow
Q4_K_Mhighgoodlow
Q5_K_Mmoderatevery goodmoderate
Q8_0lowerhighhigh

Recommendation: Find the smallest quantization that still meets your quality requirements.

Automated Benchmark

import json
import time
import requests

def benchmark(model, prompt, num_predict=100):
    url = "http://localhost:11434/api/generate"
    start = time.time()
    r = requests.post(url, json={
        "model": model,
        "prompt": prompt,
        "stream": False,
        "options": {"num_predict": num_predict}
    })
    end = time.time()
    data = r.json()
    return {
        "model": model,
        "duration": end - start,
        "tokens": data.get("eval_count", 0),
        "tps": data.get("eval_count", 0) / (end - start)
    }

results = [benchmark(m, "Erkläre KI in drei Sätzen.") for m in ["llama3.1", "qwen2.5"]]
print(json.dumps(results, indent=2))

Tips

  • Restart Ollama before each run to prevent cache interference from prior operations.
  • Use the same prompt and num_predict value across all tests.
  • Run multiple iterations and calculate the average.
  • Set temperature to 0 for reproducible results.
  • Monitor hardware temperature to avoid throttling.

Common Pitfalls

  • First run is slower: The model needs to load.
  • Varying prompt length: Affects timing measurements.
  • Streaming: Synchronous measurement becomes difficult.
  • GPU throttling: Performance degrades after sustained load.
  • Wrong quantization: Model doesn’t fit in VRAM.
  • Cache effects: Model persists in memory between runs.

Further Reading

FAQ: Ollama Benchmarks

What’s a good tokens-per-second figure? It depends on the model and hardware. For local 7B models, 10-30 tokens/s is typical.

Should I always use the highest quantization? No. If quality is acceptable for your use case, a lower quantization is faster.

How do I measure VRAM? Use nvidia-smi or rocm-smi while the model is running.

Why is the second run faster? Because the model is already loaded in memory.

Can I benchmark quality automatically? Only to a limited degree. Prepared questions and manual review work best.

Sources and Further Reading

Benchmarking is essential for finding the right Ollama models for your hardware and use cases. Measure tokens per second, TTFT, VRAM and RAM consumption, and response quality. Simple Python scripts let you systematically compare models and quantization levels. By running multiple iterations with identical prompts, you’ll get reproducible results and can deliberately choose the best balance between speed and quality.

Back to Blog
Share:

Related Posts