all the models — AI benchmark observatory
← Benchmarks

DeepSWE

DeepSWE is a software engineering agent benchmark evaluated with the mini-swe-agent harness, where each task is solved in an isolated container with no internet access. It measures an agent's ability to autonomously resolve real-world coding issues end to end.

id deepswe · max 1 · 13 models reported

#ModelScore
1GPT-5.6 Sol
OpenAI
0.73