evals.report
BenchmarksLabsCompareRun guidesIn the wild

Aider Polyglot

A coding benchmark that measures how reliably an LLM can solve and apply diff-based code edits across 225 challenging Exercism exercises spanning C++, Go, Java, JavaScript, Python, and Rust, with up to two attempts per problem.

Coding% correctHigher is better

What this benchmark measures

A coding benchmark that measures how reliably an LLM can solve and apply diff-based code edits across 225 challenging Exercism exercises spanning C++, Go, Java, JavaScript, Python, and Rust, with up to two attempts per problem.

Rows on this page are sourced from public benchmark artifacts, leaderboard exports, or source-linked model reports. Each row keeps benchmark version, source model name, and available run details attached to the score.

The metric shown here is % correct. It should be interpreted within Aider Polyglot, not compared as part of a site-wide ranking.

No composite ranking
evals.report never combines benchmarks. % correct on Aider Polyglot is its own number — don’t average it with other metrics.

Frequently asked

What is Aider Polyglot?

A coding benchmark that measures how reliably an LLM can solve and apply diff-based code edits across 225 challenging Exercism exercises spanning C++, Go, Java, JavaScript, Python, and Rust, with up to two attempts per problem. It is a coding benchmark measured by % correct.

What does % correct mean on Aider Polyglot?

Aider Polyglot reports % correct (%); higher is better. Scores are shown only within Aider Polyglot and are never averaged with other benchmarks.

What is the top reported Aider Polyglot score?

Claude Opus 4.5 has the top reported score on Aider Polyglot: 89.4% (% correct).

Why do Aider Polyglot scores differ across runs?

Harness, scaffold, reasoning effort, and prompt setup change results, so two runs of the same model can differ. evals.report keeps each score with its run context so the differences stay visible.

Does evals.report rank models across benchmarks?

No. Aider Polyglot scores are shown within their own metric; evals.report never combines benchmarks into a composite ranking or a single "best model".