Benchmark Guide
What Are AIME and MATH-500? The Two Math Benchmarks Behind LLM Rankings
Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset
When a model is described as good at math, the claim usually rests on two tests: AIME, a hard competition paper for school students, and MATH-500, a sample of competition problems across seven subjects. Here is what each one is and how to read the scores.
Models ranked on AIME, MATH-500 and GPQA Diamond.
See the math ranking →AIME
The American Invitational Mathematics Examination is sat each year by students who qualify from an earlier round. It has 15 questions in 3 hours, covering algebra, geometry, number theory, counting and probability. Every answer is a whole number between 0 and 999.
That answer format is why it suits model evaluation:
- There is no partial credit and no judging. The answer matches or it does not.
- Guessing is useless. With a thousand possible answers, chance scores are near zero.
- A new paper is written every year, so there are always recent questions a model is unlikely to have seen in training.
Results are reported for a given year's paper, or for a combined set of recent years. The id8 dataset uses the AIME figure published by Epoch AI, an independent evaluator, so every model is measured the same way.
Current top AIME scores
| # | Model | Provider | AIME 2024–25 (Epoch AI) |
|---|---|---|---|
| 1 | GPT-6 Astra | OpenAI | 100 |
| 2 | Claude Fable 5.1 | Anthropic | 100 |
| 3 | Claude Fable 5 | Anthropic | 100 |
| 4 | GPT-5.6 Sol | OpenAI | 100 |
| 5 | GPT-6 Sol | OpenAI | 100 |
| 6 | Claude Opus 5.5 | Anthropic | 100 |
| 7 | Claude Sonnet 5.5 | Anthropic | 100 |
| 8 | GPT-6.1 Sol | OpenAI | 100 |
| 9 | GPT-5.6 Terra | OpenAI | 99.7 |
| 10 | Qwen3.8 Max | Alibaba | 99.4 |
82 of the 164 models we track have an AIME score.
MATH-500
The MATH dataset was published in 2021 with 12,500 problems drawn from mathematics competitions. It spans seven subjects: prealgebra, algebra, number theory, counting and probability, geometry, intermediate algebra and precalculus, each at five difficulty levels. MATH-500 is a fixed sample of 500 of those problems that became the common test set.
Answers are short expressions, which are checked for mathematical equivalence, so 1/2 and 0.5 both count.
Current top MATH-500 scores
| # | Model | Provider | MATH-500 |
|---|---|---|---|
| 1 | DeepSeek R1 | DeepSeek | 97.3 |
| 2 | DeepSeek R1 (Full) | DeepSeek | 97.3 |
| 3 | Nemotron Ultra 253B | NVIDIA | 97 |
| 4 | o1 | OpenAI | 94.8 |
| 5 | DeepSeek R1 Distill 70B | DeepSeek | 94.5 |
| 6 | DeepSeek R1 Distill 32B | DeepSeek | 94.3 |
| 7 | DeepSeek R1 Distill 14B | DeepSeek | 93.9 |
| 8 | DeepSeek R1 Distill 7B | DeepSeek | 92.8 |
| 9 | Gemini 2.5 Pro | 92 | |
| 10 | QwQ 32B | Alibaba | 90.6 |
19 models in our dataset have a published MATH-500 score. Many recent releases no longer report it, because the leaders are near the ceiling.
Why both are used
| AIME | MATH-500 | |
|---|---|---|
| Questions | 15 per year | 500 |
| Difficulty | Hard, olympiad qualifier | From easy to hard |
| Still separates top models | Yes, though the best are near the top | No, leaders are near 100% |
| Risk of training-data leakage | Lower for the newest paper | Higher, public since 2021 |
| Margin of error | Large, since there are few questions | Small |
With only 15 questions per paper, one question changes an AIME score by nearly seven points. Treat small AIME differences with caution.
Why reasoning models do better
Models with a reasoning or thinking mode work through hidden intermediate steps before answering. On multi-step problems this lets them check and correct themselves, and it raises math scores substantially. The cost is more output tokens and longer waits.
What these benchmarks do not show
- Proofs. Both check final answers only. A correct answer reached by flawed reasoning still scores.
- Arithmetic reliability. A model that solves olympiad problems can still slip on a long multiplication. For calculation, have the model write and run code.
- Applied work. Statistics for a business report or a physics simulation involve judgement that competition problems do not test.
- Research mathematics. Newer benchmarks of unpublished research-level problems exist for this, and scores there are much lower.
How id8 uses them
Our math grade combines AIME, MATH-500 and GPQA Diamond. A missing score is estimated from the model's other results, so a model is not penalised when a vendor has not reported one of them. The full table is at Best LLM for Math.
Frequently asked questions
Which model scores highest on AIME?
As of October 2, 2026, the highest AIME result in the id8 dataset is GPT-6 Astra (100).
Which model scores highest on MATH-500?
The highest published MATH-500 score in our dataset is DeepSeek R1 (97.3).
What is the AIME?
The American Invitational Mathematics Examination: 15 questions in 3 hours, each with a whole-number answer from 0 to 999. It is the second round of the selection for the US Mathematical Olympiad team.
What is MATH-500?
A 500-problem sample of the MATH dataset, which holds 12,500 competition problems in seven subjects at five difficulty levels.
Is MATH-500 still useful?
Less than it was. Top models now score in the high nineties, so it no longer separates the leaders. It still shows whether a smaller model is competent.