Benchmark Guide
What is GPQA Diamond? A Complete Guide for 2026
Last updated: October 2, 2026
GPQA Diamond is widely considered the hardest publicly available reasoning benchmark for AI models. It was specifically designed to be "Google-proof" — questions that cannot be answered by searching the web, but require genuine expert-level understanding of biology, chemistry, and physics. When you hear about a model "surpassing PhD performance," GPQA Diamond is usually the benchmark behind that claim.
Current top scores: GPQA Diamond
Generated from the id8 dataset on October 2, 2026. Scores quoted further down this article were correct when it was written and may since have been overtaken; this table is the current picture.
| # | Model | Provider | GPQA Diamond |
|---|---|---|---|
| 1 | GPT-6 Astra | OpenAI | 95.8 |
| 2 | Claude Sonnet 5.5 | Anthropic | 95.6 |
| 3 | Gemini 3.8 Flash | 95.4 | |
| 4 | GPT-6.1 Sol | OpenAI | 95.4 |
| 5 | Gemini 3.7 Flash | 94.8 | |
| 6 | Gemini 3.1 Pro | 94.3 | |
| 7 | GPT-6 Sol | OpenAI | 94.3 |
| 8 | Claude Opus 4.7 | Anthropic | 94.2 |
See every model in the LLM Leaderboard.
What Does GPQA Diamond Measure?
GPQA (Graduate-Level Google-Proof Q&A) Diamond measures genuine expert-level reasoning in the natural sciences. The benchmark consists of 448 multiple-choice questions written by domain experts with PhDs in biology, chemistry, and physics. Each question was designed so that non-experts — even those with a solid general science background — could not reliably answer it by searching the internet. Questions require multi-step reasoning, deep domain knowledge, and the ability to distinguish between subtly different scientific concepts.
The "Diamond" subset is the hardest tier within GPQA: questions where domain experts showed the highest inter-annotator disagreement, indicating genuine difficulty rather than trivia. Source paper: arxiv.org/abs/2311.12022
Why GPQA Diamond Is a Uniquely Trusted Benchmark
Most benchmarks become contaminated quickly — questions leak into training data, and high scores stop meaning what they should. GPQA Diamond resists this for several structural reasons:
- Expert authorship — Every question was written by a subject-matter PhD, not scraped from textbooks or websites. The questions are not verbatim anywhere online.
- Google-proof design — The benchmark was specifically validated by checking that non-expert humans with internet access could not reliably answer the questions. If a question was too searchable, it was removed.
- Human calibration baselines — Non-expert humans score roughly 34% (barely above the 25% random baseline for 4-choice questions). PhD-level domain experts in the relevant field score approximately 65%. This gives a meaningful human performance ladder to compare against.
- Multi-step reasoning requirement — Correct answers require chaining multiple scientific concepts, not just recalling an isolated fact.
How GPQA Diamond Scores Work
GPQA Diamond uses accuracy (percentage of 448 questions answered correctly) with 4-choice multiple-choice format. Here is what the score range means:
- 25% — Random chance baseline; no model should score this low
- 34% — Non-expert human with internet access; models at this level have minimal science reasoning
- 50–65% — Strong general capability; approaching or matching PhD-level on some sub-topics
- 65% — PhD domain expert human performance baseline
- 70–80% — Significantly superhuman on this benchmark; frontier models from 2024–2025
- 80–90% — Far exceeds PhD performance; top frontier models from late 2025–2026
- 90%+ — Current state of the art in mid-2026
A model scoring above 90% on GPQA Diamond is, by the measure of this benchmark, significantly better than the world's best PhD researchers at answering these specific questions.
2026 Top-Scoring Models on GPQA Diamond
| Rank | Model | GPQA Diamond Score |
|---|---|---|
| 1 | Claude Opus 4.7 | 94.2% |
| 2 | GPT-5.4 | 92.8% |
| 3 | Claude Opus 4.6 | 91.3% |
Scores sourced from official model release papers and third-party evaluations. Updated June 2026.
Known Weaknesses of GPQA Diamond
Despite its reputation as one of the most rigorous benchmarks, GPQA Diamond has meaningful limitations:
- Only three domains — Biology, chemistry, and physics. GPQA Diamond tells you nothing about a model's ability in mathematics, computer science, law, history, or any other field.
- Multiple-choice only — Real scientific reasoning often requires generating explanations, writing proofs, or producing novel analysis. Selecting the correct answer from 4 choices is a much easier task.
- Small sample size — 448 questions across three domains means each domain has roughly 150 questions. Statistical noise is significant; a few lucky or unlucky guesses on ambiguous questions can shift a model's score by 1–2%.
- No practical task completion — Scoring well on GPQA Diamond does not mean a model can run an experiment, write a research paper, or synthesize a novel compound.
- Contamination risk grows over time — As GPQA questions become more widely discussed in AI research communities, future models may inadvertently see them in training data.
For broader knowledge testing, pair GPQA Diamond with MMLU-Pro. For practical capability, add SWE-bench Verified.
Primary Source
GPQA paper: GPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023)
See how all 93 models compare on GPQA Diamond — free, no signup
See Full LLM Leaderboard →Related Benchmark Guides
- What is MMLU-Pro? Expert Knowledge Benchmark Explained
- What is SWE-bench? Real Coding Benchmark Explained
- What is Chatbot Arena? Human Preference ELO Explained
- What is HumanEval? Python Coding Benchmark Explained
- What is LiveCodeBench? Anti-Contamination Coding Benchmark
- LLM Leaderboard — Compare All 93 Models Free