Benchmark Guide

What is GPQA Diamond? A Complete Guide for 2026

Last updated: October 2, 2026

GPQA Diamond is widely considered the hardest publicly available reasoning benchmark for AI models. It was specifically designed to be "Google-proof" — questions that cannot be answered by searching the web, but require genuine expert-level understanding of biology, chemistry, and physics. When you hear about a model "surpassing PhD performance," GPQA Diamond is usually the benchmark behind that claim.

Current top scores: GPQA Diamond

Generated from the id8 dataset on October 2, 2026. Scores quoted further down this article were correct when it was written and may since have been overtaken; this table is the current picture.

#ModelProviderGPQA Diamond
1GPT-6 AstraOpenAI95.8
2Claude Sonnet 5.5Anthropic95.6
3Gemini 3.8 FlashGoogle95.4
4GPT-6.1 SolOpenAI95.4
5Gemini 3.7 FlashGoogle94.8
6Gemini 3.1 ProGoogle94.3
7GPT-6 SolOpenAI94.3
8Claude Opus 4.7Anthropic94.2

See every model in the LLM Leaderboard.

What Does GPQA Diamond Measure?

GPQA (Graduate-Level Google-Proof Q&A) Diamond measures genuine expert-level reasoning in the natural sciences. The benchmark consists of 448 multiple-choice questions written by domain experts with PhDs in biology, chemistry, and physics. Each question was designed so that non-experts — even those with a solid general science background — could not reliably answer it by searching the internet. Questions require multi-step reasoning, deep domain knowledge, and the ability to distinguish between subtly different scientific concepts.

The "Diamond" subset is the hardest tier within GPQA: questions where domain experts showed the highest inter-annotator disagreement, indicating genuine difficulty rather than trivia. Source paper: arxiv.org/abs/2311.12022

Why GPQA Diamond Is a Uniquely Trusted Benchmark

Most benchmarks become contaminated quickly — questions leak into training data, and high scores stop meaning what they should. GPQA Diamond resists this for several structural reasons:

  • Expert authorship — Every question was written by a subject-matter PhD, not scraped from textbooks or websites. The questions are not verbatim anywhere online.
  • Google-proof design — The benchmark was specifically validated by checking that non-expert humans with internet access could not reliably answer the questions. If a question was too searchable, it was removed.
  • Human calibration baselines — Non-expert humans score roughly 34% (barely above the 25% random baseline for 4-choice questions). PhD-level domain experts in the relevant field score approximately 65%. This gives a meaningful human performance ladder to compare against.
  • Multi-step reasoning requirement — Correct answers require chaining multiple scientific concepts, not just recalling an isolated fact.

How GPQA Diamond Scores Work

GPQA Diamond uses accuracy (percentage of 448 questions answered correctly) with 4-choice multiple-choice format. Here is what the score range means:

  • 25% — Random chance baseline; no model should score this low
  • 34% — Non-expert human with internet access; models at this level have minimal science reasoning
  • 50–65% — Strong general capability; approaching or matching PhD-level on some sub-topics
  • 65% — PhD domain expert human performance baseline
  • 70–80% — Significantly superhuman on this benchmark; frontier models from 2024–2025
  • 80–90% — Far exceeds PhD performance; top frontier models from late 2025–2026
  • 90%+ — Current state of the art in mid-2026

A model scoring above 90% on GPQA Diamond is, by the measure of this benchmark, significantly better than the world's best PhD researchers at answering these specific questions.

2026 Top-Scoring Models on GPQA Diamond

Rank Model GPQA Diamond Score
1Claude Opus 4.794.2%
2GPT-5.492.8%
3Claude Opus 4.691.3%

Scores sourced from official model release papers and third-party evaluations. Updated June 2026.

Known Weaknesses of GPQA Diamond

Despite its reputation as one of the most rigorous benchmarks, GPQA Diamond has meaningful limitations:

  • Only three domains — Biology, chemistry, and physics. GPQA Diamond tells you nothing about a model's ability in mathematics, computer science, law, history, or any other field.
  • Multiple-choice only — Real scientific reasoning often requires generating explanations, writing proofs, or producing novel analysis. Selecting the correct answer from 4 choices is a much easier task.
  • Small sample size — 448 questions across three domains means each domain has roughly 150 questions. Statistical noise is significant; a few lucky or unlucky guesses on ambiguous questions can shift a model's score by 1–2%.
  • No practical task completion — Scoring well on GPQA Diamond does not mean a model can run an experiment, write a research paper, or synthesize a novel compound.
  • Contamination risk grows over time — As GPQA questions become more widely discussed in AI research communities, future models may inadvertently see them in training data.

For broader knowledge testing, pair GPQA Diamond with MMLU-Pro. For practical capability, add SWE-bench Verified.

Primary Source

GPQA paper: GPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023)

See how all 93 models compare on GPQA Diamond — free, no signup

See Full LLM Leaderboard →

Related Benchmark Guides