Benchmark Guide

What is SWE-bench? A Complete Guide for 2026

Last updated: October 2, 2026

SWE-bench Verified is the most respected benchmark for measuring practical software engineering ability in AI models. Unlike toy coding problems, it uses real GitHub issues from real open-source projects — the kind of messy, context-heavy bugs that actual software engineers face every day. A high SWE-bench score is the strongest available signal that a model can function as a genuine coding assistant.

Current top scores: SWE-bench Verified

Generated from the id8 dataset on October 2, 2026. Scores quoted further down this article were correct when it was written and may since have been overtaken; this table is the current picture.

#ModelProviderSWE-bench Verified
1Claude Fable 5Anthropic95
2Claude Opus 4.8Anthropic88.6
3Claude Opus 4.7Anthropic87.6
4Claude Opus 4.5Anthropic80.9
5Claude Opus 4.6Anthropic80.8
6Gemini 3.1 ProGoogle80.6
7DeepSeek V4-ProDeepSeek80.6
8o4-miniOpenAI80.5

See every model in the LLM Leaderboard.

What Does SWE-bench Verified Measure?

SWE-bench Verified measures the percentage of real GitHub issues a model can independently resolve. Each task presents the model with a failing test, a repository of production code, and a natural-language description of the bug or feature request. The model must write code changes that make the failing tests pass without breaking anything else. There are no multiple-choice options, no synthetic problems — just real repositories and real issues.

The "Verified" suffix is important: the full SWE-bench dataset contained some ambiguous or low-quality issues. The Verified subset consists of 500 manually reviewed issues where evaluators confirmed that the issue is solvable, the test suite reliably validates the fix, and the difficulty level is appropriate for benchmarking. This makes SWE-bench Verified the standard version used for model comparisons. Source: swebench.com

Why SWE-bench Matters More Than HumanEval

Most older coding benchmarks — including HumanEval — test a model's ability to write a single Python function from a docstring. That is useful but far too simple to predict real-world coding performance. SWE-bench is different in several critical ways:

  • Real repositories, not toy problems — Issues come from popular Python projects including Django, Flask, scikit-learn, requests, and matplotlib. The model must navigate large, multi-file codebases.
  • Context matters — Solving issues requires understanding how existing code is structured, not just writing new code from scratch.
  • Tests validate correctness — A fix only counts if the automated test suite passes, eliminating subjective scoring.
  • No contamination advantage — GitHub issues were opened after most models' training cutoffs, reducing the risk of models having "seen" the answer during pre-training.

For AI coding tools and developer copilots, SWE-bench is the benchmark that matters most. HumanEval tells you if a model can write a bubble sort; SWE-bench tells you if it can fix a Django ORM edge case.

How SWE-bench Scores Work

The key metric is simple: percentage of the 500 Verified issues resolved. An issue is "resolved" if and only if the model's code change passes all relevant tests in the repository's test suite. There are no partial credits — either the tests pass or they do not. Here is how to interpret the range:

  • Below 20% — Limited practical value for agentic coding tasks; only handles trivial issues
  • 20–50% — Useful coding assistant; can handle straightforward bug fixes with good prompting
  • 50–70% — Strong coding agent; resolves the majority of well-scoped issues autonomously
  • 70–85% — Frontier-level; approaches senior developer quality on Python repositories
  • 85%+ — State of the art in mid-2026; only the very best models reach this range

2026 Top-Scoring Models on SWE-bench Verified

Rank Model SWE-bench Verified Score
1Claude Opus 4.787.6%
2Claude Opus 4.680.8%
3MiniMax M2.580.2%

Scores sourced from official model release papers and SWE-bench leaderboard. Updated June 2026.

Known Weaknesses of SWE-bench

SWE-bench is the best practical coding benchmark available, but it comes with important caveats:

  • Python-only — All 12 repositories are Python projects. A model that scores well here may underperform on TypeScript, Rust, Go, or Java codebases.
  • Open-source repositories only — Real enterprise code is often more complex, more poorly documented, and more intertwined than the open-source projects used here.
  • Test quality varies — The "Verified" subset helps, but some tests are still weak validators. A model can pass a test without actually solving the bug correctly.
  • Agentic scaffolding matters — Scores depend heavily on the harness used. The same base model can score very differently depending on how the agent loop, tool use, and context window are configured.
  • 500 issues is a small sample — With only 500 tasks, a few especially hard or easy issues can shift a model's percentage point by more than it should.

Primary Source

Official SWE-bench site and leaderboard: swebench.com

See how all 93 models compare on SWE-bench Verified — free, no signup

See Full LLM Leaderboard →

Related Benchmark Guides