Benchmark Guide

What Are AIME and MATH-500? The Two Math Benchmarks Behind LLM Rankings

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

When a model is described as good at math, the claim usually rests on two tests: AIME, a hard competition paper for school students, and MATH-500, a sample of competition problems across seven subjects. Here is what each one is and how to read the scores.

Models ranked on AIME, MATH-500 and GPQA Diamond.

See the math ranking →

AIME

The American Invitational Mathematics Examination is sat each year by students who qualify from an earlier round. It has 15 questions in 3 hours, covering algebra, geometry, number theory, counting and probability. Every answer is a whole number between 0 and 999.

That answer format is why it suits model evaluation:

  • There is no partial credit and no judging. The answer matches or it does not.
  • Guessing is useless. With a thousand possible answers, chance scores are near zero.
  • A new paper is written every year, so there are always recent questions a model is unlikely to have seen in training.

Results are reported for a given year's paper, or for a combined set of recent years. The id8 dataset uses the AIME figure published by Epoch AI, an independent evaluator, so every model is measured the same way.

Current top AIME scores

#ModelProviderAIME 2024–25 (Epoch AI)
1GPT-6 AstraOpenAI100
2Claude Fable 5.1Anthropic100
3Claude Fable 5Anthropic100
4GPT-5.6 SolOpenAI100
5GPT-6 SolOpenAI100
6Claude Opus 5.5Anthropic100
7Claude Sonnet 5.5Anthropic100
8GPT-6.1 SolOpenAI100
9GPT-5.6 TerraOpenAI99.7
10Qwen3.8 MaxAlibaba99.4

82 of the 164 models we track have an AIME score.

MATH-500

The MATH dataset was published in 2021 with 12,500 problems drawn from mathematics competitions. It spans seven subjects: prealgebra, algebra, number theory, counting and probability, geometry, intermediate algebra and precalculus, each at five difficulty levels. MATH-500 is a fixed sample of 500 of those problems that became the common test set.

Answers are short expressions, which are checked for mathematical equivalence, so 1/2 and 0.5 both count.

Current top MATH-500 scores

#ModelProviderMATH-500
1DeepSeek R1DeepSeek97.3
2DeepSeek R1 (Full)DeepSeek97.3
3Nemotron Ultra 253BNVIDIA97
4o1OpenAI94.8
5DeepSeek R1 Distill 70BDeepSeek94.5
6DeepSeek R1 Distill 32BDeepSeek94.3
7DeepSeek R1 Distill 14BDeepSeek93.9
8DeepSeek R1 Distill 7BDeepSeek92.8
9Gemini 2.5 ProGoogle92
10QwQ 32BAlibaba90.6

19 models in our dataset have a published MATH-500 score. Many recent releases no longer report it, because the leaders are near the ceiling.

Why both are used

AIMEMATH-500
Questions15 per year500
DifficultyHard, olympiad qualifierFrom easy to hard
Still separates top modelsYes, though the best are near the topNo, leaders are near 100%
Risk of training-data leakageLower for the newest paperHigher, public since 2021
Margin of errorLarge, since there are few questionsSmall

With only 15 questions per paper, one question changes an AIME score by nearly seven points. Treat small AIME differences with caution.

Why reasoning models do better

Models with a reasoning or thinking mode work through hidden intermediate steps before answering. On multi-step problems this lets them check and correct themselves, and it raises math scores substantially. The cost is more output tokens and longer waits.

What these benchmarks do not show

  • Proofs. Both check final answers only. A correct answer reached by flawed reasoning still scores.
  • Arithmetic reliability. A model that solves olympiad problems can still slip on a long multiplication. For calculation, have the model write and run code.
  • Applied work. Statistics for a business report or a physics simulation involve judgement that competition problems do not test.
  • Research mathematics. Newer benchmarks of unpublished research-level problems exist for this, and scores there are much lower.

How id8 uses them

Our math grade combines AIME, MATH-500 and GPQA Diamond. A missing score is estimated from the model's other results, so a model is not penalised when a vendor has not reported one of them. The full table is at Best LLM for Math.

Frequently asked questions

Which model scores highest on AIME?

As of October 2, 2026, the highest AIME result in the id8 dataset is GPT-6 Astra (100).

Which model scores highest on MATH-500?

The highest published MATH-500 score in our dataset is DeepSeek R1 (97.3).

What is the AIME?

The American Invitational Mathematics Examination: 15 questions in 3 hours, each with a whole-number answer from 0 to 999. It is the second round of the selection for the US Mathematical Olympiad team.

What is MATH-500?

A 500-problem sample of the MATH dataset, which holds 12,500 competition problems in seven subjects at five difficulty levels.

Is MATH-500 still useful?

Less than it was. Top models now score in the high nineties, so it no longer separates the leaders. It still shows whether a smaller model is competent.

Related