Benchmark Guide

What Is Humanity's Last Exam? A Guide to the Hardest LLM Benchmark

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

Humanity's Last Exam, usually shortened to HLE, was built because models had begun to score above 90% on older tests. It asks questions that sit at the edge of expert knowledge, and scores remain far lower than on other benchmarks. This guide explains what it measures and how to read the numbers.

Models ranked on GPQA Diamond, Humanity's Last Exam and MMLU-Pro.

See the reasoning ranking →

Why it exists

A benchmark is useful while it separates models. By 2024 the strongest models were scoring above 90% on widely used tests such as MMLU, and a two-point gap on a saturated test says little. Humanity's Last Exam was designed to restore headroom: a test hard enough that progress would be visible for years.

How it was built

  • Subject experts around the world wrote questions in their own fields.
  • Each question has a single, checkable answer, either multiple choice or a short exact answer.
  • Questions were tested against leading models, and those the models already answered correctly were removed.
  • The remaining questions were reviewed by other experts.
  • The final public set has about 2,500 questions. A portion require reading an image, such as a diagram or an inscription.

A private held-out set is kept back to detect models that have simply memorised the public questions.

What it covers

Mathematics makes up the largest share. The rest spans physics, chemistry, biology and medicine, computer science, engineering, the humanities and social sciences. The questions are not trivia; many need several steps of specialist reasoning.

Current top scores

#ModelProviderHumanity's Last Exam
1GPT-6 AstraOpenAI54.8
2Claude Fable 5.1Anthropic46.5
3Gemini 3.1 ProGoogle46.4
4Gemini 3.8 FlashGoogle44.5
5Claude Opus 4.7Anthropic36.2
6Claude Opus 4.6Anthropic34.4
7Kimi K2.5Moonshot AI24.4
8Gemini 3.1 Flash LiteGoogle8.6
9Llama 4 MaverickMeta5.7

Scores are the percentage answered correctly. 9 of the 164 models we track have a published HLE score in our dataset; many vendors do not report it.

How to read a score

  • Absolute numbers are low on purpose. A score that would be poor on another benchmark can be state of the art here.
  • Compare like with like. Scores depend on whether the model could use tools such as search or code execution. A tool-assisted score is not comparable to a plain one.
  • Small gaps are noise. With about 2,500 questions, a difference of a point or two is within the margin of error.
  • Calibration is also reported. HLE asks models for a confidence level. A model that is confidently wrong is penalised on that measure, which matters for real use.

What it does not tell you

  • Usefulness for ordinary work. Writing clear emails, fixing bugs and summarising reports are not what HLE tests.
  • Open-ended research ability. Every question has a known answer. Real research does not.
  • Reliability. A model can answer a hard question and still make a careless error on an easy one.

Criticisms

  • The name overstates what a fixed question set can show.
  • Some answers have been disputed by other experts, which is common for questions at this level.
  • Like any public benchmark, it will leak into training data over time. The held-out set limits the damage but does not remove it.

How id8 uses it

HLE is one of three benchmarks behind our reasoning grade, with GPQA Diamond and MMLU-Pro. Because few models report it, a missing HLE score does not count against a model; the ranking estimates it from the model's other results.

Related benchmarks

Frequently asked questions

Which model scores highest on Humanity's Last Exam?

As of October 1, 2026, the highest published score in the id8 dataset is GPT-6 Astra (54.8).

How many questions are in Humanity's Last Exam?

About 2,500, across mathematics, the sciences, the humanities and other fields. A portion of them include an image.

Who created it?

The Center for AI Safety and Scale AI, with questions contributed by subject experts from many institutions.

Why are HLE scores so low?

Questions were kept only if leading models of the time answered them wrongly, so the test starts from the hardest material by design.

Does a high HLE score mean a model is better for my work?

Not by itself. HLE measures closed-ended expert questions. Everyday tasks such as writing, coding and summarising depend on other abilities.

Related