all the models — AI benchmark observatory
← Benchmarks

SWE-Bench Pro

SWE-Bench Pro is an advanced version of SWE-Bench that evaluates language models on complex, real-world software engineering tasks requiring extended reasoning and multi-step problem solving.

id swe-bench-pro · max 1 · 59 models reported

#ModelScore
1Claude Opus 5.5
Anthropic
0.90
2Claude Fable 5
Anthropic
0.80
3Claude Mythos Preview
Anthropic
0.78