← Benchmarks
Aider-Polyglot Edit
A challenging multi-language coding benchmark that evaluates models' code editing abilities across C++, Go, Java, JavaScript, Python, and Rust. Contains 225 of Exercism's most difficult programming problems, selected as problems that were solved by 3 or fewer out of 7 top coding models. The benchmark focuses on code editing tasks and measures both correctness of solutions and proper edit format usage. Designed to re-calibrate evaluation scales so top models score between 5-50%.
id aider-polyglot-edit · max 1 · 10 models reported
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek-V3 DeepSeek · open | 0.80 |
| 2 | Gemini 2.5 Pro Google | 0.73 |
