← Benchmarks
Aider-Polyglot
A coding benchmark that evaluates LLMs on 225 challenging Exercism programming exercises across C++, Go, Java, JavaScript, Python, and Rust. Models receive two attempts to solve each problem, with test error feedback provided after the first attempt if it fails. The benchmark measures both initial problem-solving ability and capacity to edit code based on error feedback, providing an end-to-end evaluation of code generation and editing capabilities across multiple programming languages.
id aider-polyglot · max 1 · 22 models reported
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 OpenAI | 0.88 |
| 2 | Gemini 2.5 Pro Preview 06-05 Google | 0.82 |
| 3 | o3 OpenAI | 0.81 |
| 4 | Gemini 2.5 Pro Google | 0.77 |
| 5 | DeepSeek-V3.2-Exp DeepSeek · open | 0.74 |
| 6 | DeepSeek-R1-0528 DeepSeek · open | 0.72 |
