Benchmark Guide

What is HumanEval? A Complete Guide for 2026

Last updated: October 2, 2026

HumanEval was one of the first serious benchmarks for measuring AI coding ability and it remains one of the most widely cited numbers in model release announcements. But by 2026, it has a significant problem: it is nearly saturated. Multiple models now score above 95%, and some approach 99%. Understanding what HumanEval measures — and more importantly, what it does not measure — is essential for reading AI benchmark claims critically.

Current top scores: HumanEval

Generated from the id8 dataset on October 2, 2026. Scores quoted further down this article were correct when it was written and may since have been overtaken; this table is the current picture.

#ModelProviderHumanEval
1Kimi K2.5Moonshot AI99
2Claude Sonnet 4.5Anthropic97.6
3GPT-5OpenAI96.7
4Qwen 3 32BAlibaba95.2
5Claude Opus 4.6Anthropic95
6Gemini 2.5 ProGoogle93
7Mistral Small 3.2 24BMistral AI92.9
8Qwen 2.5-Coder 32BAlibaba92.7

See every model in the LLM Leaderboard.

What Does HumanEval Measure?

HumanEval measures a language model's ability to write correct Python functions from natural language descriptions. The benchmark was released by OpenAI in 2021 and consists of 164 hand-crafted programming problems. Each problem provides a function signature, a docstring describing what the function should do, and a set of example inputs and outputs. The model must complete the function body.

Correctness is measured by running each candidate solution against a hidden set of unit tests. A solution is considered correct only if it passes all tests — no partial credit. The primary metric is pass@1: the probability that a single model generation passes all tests. Source: github.com/openai/human-eval

Why HumanEval Was Important

When HumanEval was released in 2021, the state of the art was around 28% pass@1 (with OpenAI's Codex). By 2023, models were reaching 80%. By 2024, 90%. By mid-2026, multiple models score above 95% and at least one approaches 99%. HumanEval was crucial for its time because it established a standard, reproducible, automated way to evaluate code generation — something that previously required expensive human review of model outputs.

Its simplicity is also part of why it became a standard: the problems are clearly specified, the test harness is publicly available, and the metric (pass@1) is easy to understand. These properties made it possible for every lab to run the same evaluation and compare results.

How HumanEval Scores Work: What Is pass@1?

The pass@1 metric measures whether the model's first (and only) attempt at writing the function passes all unit tests. This is distinct from pass@k metrics (like pass@10 or pass@100), which measure whether at least one of k independently sampled generations passes. Here is how to interpret the 2026 score range:

  • Below 50% — Weak coding ability; only handles the simplest problems
  • 50–75% — Reasonable for basic code generation; struggles with complex logic
  • 75–90% — Good general coding; the threshold for "useful coding assistant" before 2024
  • 90–97% — Strong but no longer distinctive in 2026; many frontier models cluster here
  • 97%+ — Approaching the practical ceiling; meaningful differences are hard to detect

A HumanEval score below 80% in 2026 is a meaningful red flag for a model marketed as a "coding model." But differences between 90% and 96% tell you very little about real-world coding performance — the problems at the margin between these scores are borderline trivial.

2026 Top-Scoring Models on HumanEval

Rank Model HumanEval pass@1
1Kimi K2.599%
2Claude Opus 4.695%
3Grok 394.5%

Scores sourced from official model release papers and third-party evaluations. Updated June 2026. Note: high clustering near the ceiling makes rankings less meaningful than on harder benchmarks.

Known Weaknesses of HumanEval

HumanEval has significant limitations that become more important the higher scores climb:

  • Saturated — limited discrimination at the frontier — When the best models score 95–99%, the benchmark can no longer meaningfully differentiate between them. A 1–2% difference on 164 problems is often within statistical noise.
  • Small test set — 164 problems is a tiny sample. A model can have a good day (or a lucky prompt) and score a few points higher than its true capability. The variance is significant.
  • Heavy memorization risk — HumanEval has been public since 2021. Many training datasets include the exact function signatures, docstrings, and expected outputs. Models may be recalling solutions rather than generating them. LiveCodeBench was specifically designed to address this problem.
  • Only Python, only isolated functions — Real coding involves multi-file projects, refactoring existing code, debugging, and cross-language work. Writing a single Python function from a docstring is a thin slice of what a useful coding assistant needs to do.
  • Weak test coverage — Some problems have test suites that are too simple to catch bugs. A model can pass the tests with an incorrect algorithm that happens to work on the provided cases.
  • No context of an existing codebase — HumanEval problems are entirely self-contained. Real code lives in a codebase with conventions, dependencies, and architectural constraints that the model must understand and respect.

For a more realistic coding evaluation in 2026, use SWE-bench Verified for practical software engineering and LiveCodeBench for contamination-resistant algorithmic problems.

Primary Source

HumanEval repository: github.com/openai/human-eval

See how all 93 models compare on HumanEval — free, no signup

See Full LLM Leaderboard →

Related Benchmark Guides