all the models — AI benchmark observatory
← Benchmarks

Aider-Polyglot

A coding benchmark that evaluates LLMs on 225 challenging Exercism programming exercises across C++, Go, Java, JavaScript, Python, and Rust. Models receive two attempts to solve each problem, with test error feedback provided after the first attempt if it fails. The benchmark measures both initial problem-solving ability and capacity to edit code based on error feedback, providing an end-to-end evaluation of code generation and editing capabilities across multiple programming languages.

id aider-polyglot · max 1 · 22 models reported

#ModelScore
1GPT-5
OpenAI
0.88
2Gemini 2.5 Pro Preview 06-05
Google
0.82
3o3
OpenAI
0.81
4Gemini 2.5 Pro
Google
0.77
5DeepSeek-V3.2-Exp
DeepSeek · open
0.74
6DeepSeek-R1-0528
DeepSeek · open
0.72