← Benchmarks
DeepSWE
DeepSWE is a software engineering agent benchmark evaluated with the mini-swe-agent harness, where each task is solved in an isolated container with no internet access. It measures an agent's ability to autonomously resolve real-world coding issues end to end.
id deepswe · max 1 · 13 models reported
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol OpenAI | 0.73 |
