Benchmark Guide

What Is the Aider Polyglot Benchmark? Code Editing Across Six Languages

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

Writing new code and editing existing code are different skills. Aider Polyglot measures the second: given a file and a failing test, can the model change the file so the tests pass, in a format a tool can apply? It covers six languages, which makes it a useful complement to Python-only tests.

Models ranked on SWE-bench Verified, Arena Elo (WebDev) and Aider Polyglot.

See the agents ranking →

What the benchmark asks

Each of the 225 problems is a programming exercise with a starter file and a set of unit tests. The model is asked to edit the file so that the tests pass.

  1. The model receives the instructions and the starter code.
  2. It replies with an edit.
  3. The tests are run.
  4. If they fail, the model is shown the test output and gets one more attempt.

The headline number is the percentage of exercises solved within those two attempts.

The part that makes it different: edit format

A coding assistant cannot paste a model's reply into your project by hand. The model has to express its change in a structured form, such as a search-and-replace block or a diff, that the tool can apply to the file. If the block does not match the file exactly, the edit fails even when the idea was right.

Aider Polyglot reports how often the model produced well-formed edits, alongside how often the code was correct. A model with strong reasoning and sloppy edit formatting scores worse here than on tests that only look at the final code. For anyone using an AI coding tool, that is the behaviour that matters day to day.

Why six languages

Many coding benchmarks are Python only. Python is the language models have seen most, so Python scores flatter them. The exercises here are spread across C++, Go, Java, JavaScript, Python and Rust, which shows whether ability carries over to languages with stricter compilers and less training data.

The 225 exercises were selected as the hardest from a much larger pool, after an earlier Python-only version of the benchmark became too easy for leading models.

Current top scores

#ModelProviderAider Polyglot
1Qwen 3 235B-A22BAlibaba59.6
2DeepSeek R1DeepSeek56.9
3DeepSeek V3DeepSeek48.4
4GPT-oss 120BOpenAI41.8
5Qwen 3 32BAlibaba40
6QwQ 32BAlibaba20.9
7Qwen 2.5-Coder 32BAlibaba16.4
8Llama 4 MaverickMeta15.6
9Gemma 3 27BGoogle4.9

9 of the 164 models we track have a published Aider Polyglot score in our dataset. Results depend on settings such as reasoning effort and edit format, so compare runs with similar settings.

Aider Polyglot and SWE-bench compared

Aider PolyglotSWE-bench Verified
Tasks225 self-contained exercises500 real GitHub issues
LanguagesSixPython
Codebase sizeOne or a few filesLarge repositories
Tests navigation of a codebaseNoYes
Tests clean, applicable editsYes, explicitlyIndirectly
AttemptsTwoDepends on the agent

They answer different questions. SWE-bench asks whether a model can find and fix a problem in a big project. Aider Polyglot asks whether it can make a precise change in several languages. A good coding model does well on both. See What is SWE-bench?.

Limits

  • Exercises, not projects. Practice problems have clear specifications and complete tests. Real work rarely does.
  • Public problems. The exercises are on the open web and are likely in training data.
  • One tool's view. The benchmark is run through Aider, with its prompts and edit formats. Another tool may suit a given model better or worse.
  • Cost is not in the score. The leaderboard published by Aider lists the cost of each run; a model that scores slightly higher may cost many times more.

How id8 uses it

Aider Polyglot is one of the three benchmarks behind our agents grade and contributes to tool calling. Because few models report it, a missing score is estimated from the model's other coding results. See Best LLM for Agents.

Frequently asked questions

Which model scores highest on Aider Polyglot?

As of October 1, 2026, the highest published Aider Polyglot score in the id8 dataset is Qwen 3 235B-A22B (59.6).

Which languages does Aider Polyglot cover?

C++, Go, Java, JavaScript, Python and Rust.

How many problems are in it?

225 exercises, chosen as the hardest from the Exercism practice sets in those six languages.

How is it different from SWE-bench?

SWE-bench uses real issues from large Python repositories and tests navigating a codebase. Aider Polyglot uses self-contained exercises in six languages and tests producing correct, cleanly applicable edits.

What is Aider?

An open-source command-line tool for pair programming with an LLM. Its authors maintain the benchmark to see which models work best in the tool.

Related