Benchmark Guide
What is HumanEval? A Complete Guide for 2026
Last updated: October 2, 2026
HumanEval was one of the first serious benchmarks for measuring AI coding ability and it remains one of the most widely cited numbers in model release announcements. But by 2026, it has a significant problem: it is nearly saturated. Multiple models now score above 95%, and some approach 99%. Understanding what HumanEval measures — and more importantly, what it does not measure — is essential for reading AI benchmark claims critically.
Current top scores: HumanEval
Generated from the id8 dataset on October 2, 2026. Scores quoted further down this article were correct when it was written and may since have been overtaken; this table is the current picture.
| # | Model | Provider | HumanEval |
|---|---|---|---|
| 1 | Kimi K2.5 | Moonshot AI | 99 |
| 2 | Claude Sonnet 4.5 | Anthropic | 97.6 |
| 3 | GPT-5 | OpenAI | 96.7 |
| 4 | Qwen 3 32B | Alibaba | 95.2 |
| 5 | Claude Opus 4.6 | Anthropic | 95 |
| 6 | Gemini 2.5 Pro | 93 | |
| 7 | Mistral Small 3.2 24B | Mistral AI | 92.9 |
| 8 | Qwen 2.5-Coder 32B | Alibaba | 92.7 |
See every model in the LLM Leaderboard.
What Does HumanEval Measure?
HumanEval measures a language model's ability to write correct Python functions from natural language descriptions. The benchmark was released by OpenAI in 2021 and consists of 164 hand-crafted programming problems. Each problem provides a function signature, a docstring describing what the function should do, and a set of example inputs and outputs. The model must complete the function body.
Correctness is measured by running each candidate solution against a hidden set of unit tests. A solution is considered correct only if it passes all tests — no partial credit. The primary metric is pass@1: the probability that a single model generation passes all tests. Source: github.com/openai/human-eval
Why HumanEval Was Important
When HumanEval was released in 2021, the state of the art was around 28% pass@1 (with OpenAI's Codex). By 2023, models were reaching 80%. By 2024, 90%. By mid-2026, multiple models score above 95% and at least one approaches 99%. HumanEval was crucial for its time because it established a standard, reproducible, automated way to evaluate code generation — something that previously required expensive human review of model outputs.
Its simplicity is also part of why it became a standard: the problems are clearly specified, the test harness is publicly available, and the metric (pass@1) is easy to understand. These properties made it possible for every lab to run the same evaluation and compare results.
How HumanEval Scores Work: What Is pass@1?
The pass@1 metric measures whether the model's first (and only) attempt at writing the function passes all unit tests. This is distinct from pass@k metrics (like pass@10 or pass@100), which measure whether at least one of k independently sampled generations passes. Here is how to interpret the 2026 score range:
- Below 50% — Weak coding ability; only handles the simplest problems
- 50–75% — Reasonable for basic code generation; struggles with complex logic
- 75–90% — Good general coding; the threshold for "useful coding assistant" before 2024
- 90–97% — Strong but no longer distinctive in 2026; many frontier models cluster here
- 97%+ — Approaching the practical ceiling; meaningful differences are hard to detect
A HumanEval score below 80% in 2026 is a meaningful red flag for a model marketed as a "coding model." But differences between 90% and 96% tell you very little about real-world coding performance — the problems at the margin between these scores are borderline trivial.
2026 Top-Scoring Models on HumanEval
| Rank | Model | HumanEval pass@1 |
|---|---|---|
| 1 | Kimi K2.5 | 99% |
| 2 | Claude Opus 4.6 | 95% |
| 3 | Grok 3 | 94.5% |
Scores sourced from official model release papers and third-party evaluations. Updated June 2026. Note: high clustering near the ceiling makes rankings less meaningful than on harder benchmarks.
Known Weaknesses of HumanEval
HumanEval has significant limitations that become more important the higher scores climb:
- Saturated — limited discrimination at the frontier — When the best models score 95–99%, the benchmark can no longer meaningfully differentiate between them. A 1–2% difference on 164 problems is often within statistical noise.
- Small test set — 164 problems is a tiny sample. A model can have a good day (or a lucky prompt) and score a few points higher than its true capability. The variance is significant.
- Heavy memorization risk — HumanEval has been public since 2021. Many training datasets include the exact function signatures, docstrings, and expected outputs. Models may be recalling solutions rather than generating them. LiveCodeBench was specifically designed to address this problem.
- Only Python, only isolated functions — Real coding involves multi-file projects, refactoring existing code, debugging, and cross-language work. Writing a single Python function from a docstring is a thin slice of what a useful coding assistant needs to do.
- Weak test coverage — Some problems have test suites that are too simple to catch bugs. A model can pass the tests with an incorrect algorithm that happens to work on the provided cases.
- No context of an existing codebase — HumanEval problems are entirely self-contained. Real code lives in a codebase with conventions, dependencies, and architectural constraints that the model must understand and respect.
For a more realistic coding evaluation in 2026, use SWE-bench Verified for practical software engineering and LiveCodeBench for contamination-resistant algorithmic problems.
Primary Source
HumanEval repository: github.com/openai/human-eval
See how all 93 models compare on HumanEval — free, no signup
See Full LLM Leaderboard →Related Benchmark Guides
- What is SWE-bench? Real Coding Benchmark Explained
- What is LiveCodeBench? Anti-Contamination Coding Benchmark
- What is GPQA Diamond? PhD-Level Reasoning Explained
- What is MMLU-Pro? Expert Knowledge Benchmark Explained
- What is Chatbot Arena? Human Preference ELO Explained
- LLM Leaderboard — Compare All 93 Models Free