Benchmark Guide

What is MMLU-Pro? A Complete Guide for 2026

Last updated: October 2, 2026

MMLU-Pro is one of the most widely cited benchmarks for measuring broad, expert-level knowledge in large language models. If you have ever seen a model advertise a score like "90.2% on MMLU-Pro," this guide explains exactly what that number means, how the test works, and what it reveals — and does not reveal — about a model's real capabilities.

Current top scores: MMLU-Pro

Generated from the id8 dataset on October 2, 2026. Scores quoted further down this article were correct when it was written and may since have been overtaken; this table is the current picture.

#ModelProviderMMLU-Pro
1o3OpenAI92.3
2Claude Opus 4.8Anthropic91.8
3Claude Fable 5Anthropic91.5
4Gemini 3.1 ProGoogle90.99
5GPT-4.1OpenAI90.2
6GPT-oss 120BOpenAI90
7Claude Opus 4.5Anthropic90
8Grok 4.3xAI88

See every model in the LLM Leaderboard.

What Does MMLU-Pro Measure?

MMLU-Pro measures breadth of factual knowledge and reasoning across 14 academic subjects, including STEM fields (biology, chemistry, physics, mathematics, engineering, computer science), professional domains (law, business, economics), and humanities (history, philosophy, psychology). The benchmark contains approximately 12,000 questions, each with 10 answer choices instead of the original MMLU's 4 choices. This makes random guessing far less rewarding and forces models to demonstrate genuine knowledge rather than process-of-elimination reasoning.

The benchmark was published by TIGER-Lab in 2024 (source paper: arxiv.org/abs/2406.01574) specifically because the original MMLU had become too easy for frontier models — many were scoring above 85–90% on it, making differentiation between top models nearly impossible.

How Is MMLU-Pro Different from Regular MMLU?

The original Massive Multitask Language Understanding (MMLU) benchmark from 2021 was groundbreaking for its time: 57 subjects, 4-choice questions, testing general academic knowledge. But within three years it was essentially saturated — frontier models were scoring above 90% and human expert performance on the hardest subjects was being matched or exceeded. MMLU-Pro addresses this by:

  • Expanding answer choices from 4 to 10 — reducing the chance of guessing correctly from 25% to 10%, which sharpens discrimination between models
  • Selecting harder questions — questions where GPT-4 itself struggled were preferentially included, raising the difficulty ceiling
  • Adding reasoning-heavy sub-topics — more multi-step problems, fewer pure recall questions
  • Removing trivially easy questions — filtering out items where essentially all capable models got perfect marks

As a result, MMLU-Pro scores are typically 10–20 percentage points lower than MMLU scores for the same model, giving much better signal at the frontier.

How Scores Work: What Is a Good MMLU-Pro Score?

MMLU-Pro uses simple accuracy: the percentage of 12,000 questions answered correctly in a single-pass, no-chain-of-thought setting (though models are generally allowed to "think" before answering). Here is how to interpret the scale:

  • Below 60% — Small or older models; weaker than a well-prepared undergraduate on most subjects
  • 60–75% — Mid-tier capable models; strong on common topics, weak on niche professional domains
  • 75–85% — Frontier-adjacent; matches or exceeds graduate-level performance on most subjects
  • 85%+ — Top frontier performance; only a handful of the largest models reach this range in 2026
  • 90%+ — Currently the cutting edge as of mid-2026

Human expert performance on the full MMLU-Pro suite sits around 55–65% for non-domain specialists and higher for subject matter experts in their specific field, making scores above 85% genuinely impressive.

2026 Top-Scoring Models on MMLU-Pro

Rank Model MMLU-Pro Score
1GPT-4.190.2%
2GPT-oss 120B90.0%
3Qwen 3.587.8%

Scores sourced from official model release papers and third-party evaluations. Updated June 2026.

Known Weaknesses of MMLU-Pro

MMLU-Pro is a strong benchmark but it has real limitations that practitioners should understand before making model selection decisions based solely on this score:

  • Memorization vs. reasoning — A significant portion of questions reward factual recall rather than reasoning. A model trained on more of the internet may score higher than a model with better reasoning but less data coverage.
  • Static questions invite data contamination — The fixed question set has been public since 2024. Models trained on or after the release may have seen these questions in pre-training data, inflating scores.
  • No open-ended generation — MMLU-Pro only measures multiple-choice ability. A model that excels here may still write mediocre code, produce incoherent long-form text, or fail on agentic tasks.
  • English-only — MMLU-Pro does not evaluate multilingual capability; a model's score here tells you nothing about Hindi, Chinese, or Arabic performance.
  • No practical task completion — Answering academic questions is not the same as resolving GitHub issues, writing production code, or managing multi-step workflows.

For a more complete picture, pair MMLU-Pro with SWE-bench Verified for coding tasks and GPQA Diamond for reasoning.

Primary Source

MMLU-Pro paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark (TIGER-Lab, 2024)

See how all 93 models compare on MMLU-Pro — free, no signup

See Full LLM Leaderboard →

Related Benchmark Guides