Skip to content
BotServBotServ
BenchmarksMMLUHumanEvalGSM8KMT-BenchComparisonModels

AI Benchmarks: Compare Models Objectively

AI benchmarks: MMLU, HumanEval, GSM8K, MT-Bench and more. How to compare models and which metrics matter.

S

schutzgeist

12 min read
AI Benchmarks: Compare Models Objectively

AI Benchmarks: Comparing Models Objectively

What This Article Covers

  • What AI benchmarks are and why they matter for model comparison
  • The key benchmarks: MMLU, HumanEval, GSM8K, MT-Bench, and LMSYS Arena
  • How to interpret and compare benchmark results correctly
  • Which benchmarks are relevant for your use case
  • What limitations and pitfalls benchmarks have

Introduction

If you’re interested in local AI, you’ll quickly ask the central question: which model is best? The answer isn’t straightforward, because “best” depends on what you want to do. A model that writes good code might struggle with German text. A model that excels at math might disappoint on creative tasks.

That’s where benchmarks come in. A benchmark is a standardized test that runs a model against a specific task and assigns a score to the result. Instead of relying on subjective impressions, a benchmark measures how well a model performs on a clearly defined task. This makes it possible to compare different models objectively.

In this article, you’ll learn the most important benchmarks, understand what they measure, and how to use the results to choose your model. If you’re still uncertain which model fits your needs, the article on finding models will help you with the initial selection.

Why Do You Need Benchmarks?

Imagine reading online that a new model is “groundbreaking.” How do you verify this? Without benchmarks, you’d be left testing the model yourself and trusting your gut. The problem: your own tests are subjective, cover only a few tasks, and are hard to compare with other models.

Benchmarks solve this problem. They ask every model the same questions and grade the answers by fixed rules. That way you can see that Model A solves math problems correctly 85 percent of the time, while Model B only manages 60 percent. This is objective, reproducible, and comparable.

For local AI use, benchmarks are especially valuable because you can’t download and test every model yourself. A benchmark result tells you whether it’s worth the download before you transfer gigabytes of model weights to your hard drive. That saves time and storage space.

Benchmarks also help you decide between model sizes. If you see a 9B model performs nearly as well as a 27B model on MMLU, you can choose the smaller one and save VRAM.

Benchmarks Explained

A benchmark consists of three parts: a dataset of questions or tasks, an evaluation method, and a score that summarizes the result. The model receives the questions, provides answers, and the evaluation method checks how many are correct.

There are two main types of benchmarks:

Multiple-choice benchmarks like MMLU give the model a question and several answer options. The model picks one, and the evaluation checks if it’s the right choice. This is simple, objective, and automatically scorable.

Open-ended benchmarks like MT-Bench let the model provide free-form answers. Evaluation is done either by human raters or by another AI model acting as a judge. This is more realistic but less objective.

Most major benchmarks use a mix of both approaches. In the next section, we’ll walk through individual benchmarks.

Who Is This Article For?

This article is for you if you:

  • Want to compare models without testing each one yourself
  • Need to understand benchmark results you see in articles and forums
  • Want to make an informed purchase or download decision
  • Need to know which benchmarks matter for your use case
  • Wonder why a model performs well on one benchmark but poorly on another

You should have a basic understanding of what an LLM is and how models work. If you’re just starting out, first read What is local AI?.

Key Terms

TermDefinition
BenchmarkStandardized test for comparing model capabilities
ScoreNumerical value summarizing benchmark results
MMLUMassive Multitask Language Understanding, broad knowledge test
HumanEvalBenchmark for programming tasks in Python
GSM8KGrade School Math, elementary-level math problems
MT-BenchMulti-Turn Benchmark for conversation quality
LMSYS ArenaPlatform for blind human comparisons between models
Few-ShotTesting method where examples precede the task
Zero-ShotTesting method where the model solves tasks without examples
LeaderboardRanking list sorting models by benchmark results

Major Benchmarks Explained

MMLU: Testing Broad Knowledge

MMLU (Massive Multitask Language Understanding) is the most well-known benchmark for general knowledge. It consists of over 15,000 multiple-choice questions across 57 subjects, from history and law to medicine and mathematics. The model must select the correct answer from four options.

MMLU measures how much knowledge a model has and how well it retrieves that knowledge across different areas. A score of 80 percent means the model answered 80 percent of questions correctly. Current top models exceed 85 percent.

For local AI, MMLU is a good initial filter. If a model scores poorly on MMLU, it probably won’t be a solid all-rounder. However, MMLU only measures knowledge, not the ability to write fluent text or solve complex problems.

HumanEval: Evaluating Programming

HumanEval is the standard benchmark for programming ability. It contains 164 programming tasks in Python. The model receives a function signature and a description and must write the code that solves the task. Evaluation is automatic, running the generated code against test cases.

A score of 70 percent means the model correctly solved 70 percent of tasks. The score is reported as “pass@1,” meaning the model gets one attempt. There’s also “pass@10,” where the model gets ten attempts and one must be correct.

For users seeking a model for programming, HumanEval is the most important benchmark. Models like DeepSeek-Coder and Qwen excel here. You can find more about these models in our articles on Llama models and Qwen models.

GSM8K: Testing Mathematical Reasoning

GSM8K (Grade School Math 8K) tests mathematical reasoning at the elementary level. The benchmark consists of over 8,000 math problems requiring only grade-school knowledge. Problems involve multiple steps and logical thinking but no advanced mathematics.

Despite the elementary level, GSM8K is surprisingly difficult for many models. The benchmark tests whether a model can think through problems step by step logically, not just whether it memorized facts. Models with reasoning capabilities like DeepSeek-R1 perform particularly well here.

The score indicates what percentage of problems the model solved correctly. Current top models exceed 90 percent, while smaller models often fall between 50 and 70 percent.

MT-Bench: Evaluating Conversation Quality

MT-Bench (Multi-Turn Benchmark) measures how well a model performs in extended conversations. The benchmark consists of predefined multi-round conversations where the model must do more than answer a single question. It needs to handle follow-up questions and maintain context throughout the exchange.

Evaluation is performed by a strong AI model acting as a judge, rating responses on a scale from 1 to 10. This is less objective than multiple-choice scoring, but it better reflects real-world usage.

MT-Bench is especially valuable when you’re looking for a model to power chat applications. A model might score well on MMLU but poorly on MT-Bench if its responses are stilted or context-deaf.

LMSYS Chatbot Arena: Human Comparison at Scale

The LMSYS Chatbot Arena is a platform where human users pit two models against each other in blind comparison. Users ask a question, see both models’ responses without knowing which is which, and vote for the better answer. From thousands of such comparisons, a ranking emerges based on genuine human preference.

The Arena uses an Elo rating system, the same one used in chess. Models that win frequently climb the rankings. The result is a leaderboard reflecting how humans actually perceive model quality.

For local AI, the Arena is particularly valuable because it measures real user questions and genuine preferences, not just standardized test tasks. That said, most compared models are large cloud-based ones, and not all local models appear on the Arena.

The table below shows typical benchmark results for popular models suited to local deployment. Values are rounded and may vary slightly depending on quantization and testing methodology.

ModelMMLUHumanEvalGSM8KMT-BenchSize
Llama 3 8Bca. 68ca. 62ca. 52ca. 8.08B
Llama 3 70Bca. 80ca. 75ca. 82ca. 8.970B
Qwen 2.5 7Bca. 74ca. 70ca. 75ca. 8.37B
Qwen 2.5 32Bca. 83ca. 82ca. 85ca. 9.032B
Gemma 2 9Bca. 72ca. 60ca. 68ca. 8.39B
Gemma 2 27Bca. 76ca. 65ca. 78ca. 8.727B
Mistral 7Bca. 62ca. 40ca. 52ca. 7.67B
DeepSeek-R1 7B Distillca. 70ca. 72ca. 80ca. 8.27B
DeepSeek-R1 32B Distillca. 79ca. 80ca. 88ca. 8.832B

A few patterns stand out. Qwen models excel at coding and math. DeepSeek-R1 Distills punch well above their weight on GSM8K. Gemma 2 27B shows surprising strength on MMLU for its size. Mistral 7B is a solid generalist but hits limitations with coding and math tasks.

For more detail on individual model families, see the articles on Llama models, Qwen models, and Mistral models.

How to Interpret Benchmark Results

Benchmark results are useful, but you need to read them correctly. Here are the key principles:

Match the benchmark to your use case. Looking for a coding model? Skip MMLU and focus on HumanEval. Need math? Check GSM8K. No single benchmark tells you whether a model will work for your specific task.

Compare models of similar size. A 7B model with 70 percent MMLU is impressive. A 70B model with 70 percent MMLU is disappointing. Always compare models in the same weight class.

Account for quantization. Benchmark results are typically obtained from unquantized models (FP16). A quantized model in Q4_K_M may score a few points lower. Budget for roughly 1 to 3 percentage points of difference with Q4_K_M.

Look for clear gaps, not decimal places. A model scoring 71.3 percent MMLU is not meaningfully better than one at 70.8 percent. Differences under 2 percentage points are often noise. Focus on substantial gaps.

Check the date. AI models evolve rapidly. A benchmark result from a year ago may be outdated. Seek the latest findings.

Running Benchmarks Yourself

If you’re running a model locally and want to test it yourself, you can use tools like lm-evaluation-harness. This is the standard tool used by most model developers.

# Install lm-evaluation-harness
pip install lm-eval

# Run MMLU with a local model
lm_eval --model hf --model_args pretrained=meta-llama/Llama-3-8B-Instruct --tasks mmlu --batch_size 8

Running benchmarks takes time and compute. For most people, comparing published results is sufficient rather than running benchmarks yourself. But if you need precise figures for a specific quantization, running them locally is the most reliable approach.

For more on quality testing of local systems, see Testing Quality.

Common Pitfalls

1. Treating one benchmark as gospel. No single benchmark captures all model abilities. A model can excel at MMLU and disappoint in conversation. Always look at multiple benchmarks.

2. Equating benchmark scores with real-world performance. A high MMLU score doesn’t mean the model will work well in your application. Benchmarks are simplified tests that only approximate real usage.

3. Ignoring quantization effects. Most benchmark results come from unquantized models. A Q4_K_M model can perform noticeably worse. Always compare at the same quantization level.

4. Overlooking size differences. A 70B model should outperform a 7B model. The interesting comparison is how much each model achieves per parameter, not absolute scores.

5. Relying on stale results. AI moves fast. A six-month-old benchmark result may already be superseded. Look for current comparisons and leaderboards.

6. Using LMSYS Arena as your only source. The Arena measures human preference, not objective capability. A model that sounds eloquent but is wrong can rank higher than one that’s terse but accurate.

7. Overlooking contamination. Some models have seen benchmark questions during training. This skews results upward. Reputable developers disclose when contamination has been ruled out.

Hardware, Cost, and Security

Hardware. Running benchmarks yourself requires compute. A full MMLU run with 15,000 questions can take hours on a GPU. For most users, reading published results is sufficient. If you do test locally, use a GPU with at least 12 GB VRAM.

Cost. Benchmark results are free. Major leaderboards like LMSYS Arena and Hugging Face Open LLM Leaderboard are publicly accessible. Running them yourself costs only compute time on your hardware.

Security. Benchmarks have no direct security implications. When you run benchmarks locally, your data stays on your machine. Nothing is sent to external services.

Further Reading

FAQ

What’s a good MMLU score?

A score above 70 percent is solid for a 7B to 13B model. Models at 30B and larger should exceed 75 percent. Top performers reach above 85 percent. Always compare models of similar size.

Are benchmarks more reliable than my own tests?

Benchmarks offer objectivity and comparability, while your own tests reflect actual usage patterns. The best approach combines both: use benchmarks for initial screening, then evaluate your top candidates on tasks that matter to you.

What does pass@1 mean in HumanEval?

pass@1 means the model has one attempt to solve the task. If the generated code passes all test cases, it counts as success. pass@10 means the model gets ten attempts, and at least one must succeed.

How reliable is the LMSYS Chatbot Arena?

The Arena draws from thousands of human comparisons and ranks among the most dependable methods for measuring model quality. However, it isn’t perfect: people often prefer longer or more conversational responses, even when they aren’t more accurate.

Does quantization affect benchmark results?

Yes. A Q4_K_M model can score 1 to 3 percentage points lower than the FP16 original. Very aggressive quantization like Q2 or Q3 can produce significantly larger gaps.

What is benchmark contamination?

Contamination occurs when a model has already seen benchmark questions during training. This inflates scores artificially because the model recognizes the answers rather than solving the task. Reputable developers screen for this.

Can I run benchmarks against models served by Ollama?

Yes, tools like lm-evaluation-harness let you run benchmarks on models served through Ollama. You query the model via the Ollama API. Execution time ranges from minutes to hours depending on the benchmark and model size.

Which benchmark matters most for developers?

HumanEval is the standard for code generation ability. It measures whether the model produces correct Python code. For more comprehensive evaluation, MBPP and LiveCodeBench offer similar tasks with different emphases.

Why do some models perform so differently on GSM8K?

GSM8K tests logical reasoning and multi-step arithmetic. Models trained on reasoning, like DeepSeek-R1, excel here because they decompose problems into intermediate steps. Models without reasoning often produce quick but incorrect answers.

Are there benchmarks for German?

Yes, multilingual benchmarks like MGSM (multilingual mathematics) and multilingual MMLU variants exist. Most standard benchmarks remain English-focused. For German-heavy applications, supplement benchmarks with your own testing.

Should I always pick the model with the highest score?

Not necessarily. A model with a slightly lower score may suit your specific task better, especially if it’s smaller and therefore faster. Benchmarks provide guidance, not the final word.

Sources

  • Hugging Face Open LLM Leaderboard
  • LMSYS Chatbot Arena Leaderboard
  • MMLU: Massive Multitask Language Understanding (Paper on arXiv)
  • HumanEval: Evaluating Code Synthesis (Paper on arXiv)
  • GSM8K: Training Verifiers for Math Word Problems (Paper on arXiv)
  • MT-Bench: Multi-Turn Benchmark (LMSYS Research)
  • lm-evaluation-harness Documentation on GitHub
Back to Blog
Share:

Related Posts