← Benchmarks
FrontierSWE
FrontierSWE measures whether an agent can complete open-ended technical projects at the scale of hours to tens of hours, spanning systems optimization, large-scale code construction, and applied ML research. Performance is reported as a dominance score, where higher is better.
id frontierswe · max 1 · 16 models reported
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5 Anthropic | 0.90 |
| 2 | Kimi K3 Moonshot AI · open | 0.81 |
| 3 | GLM-5.3 Zhipu AI · open | 0.78 |
| 4 | Claude Opus 4.8 Anthropic | 0.75 |
| 5 | GLM-5.2 Zhipu AI · open | 0.74 |
| 6 | Qwen3.8 Max Alibaba Cloud / Qwen Team · open | 0.73 |
| 7 | GPT-5.5 OpenAI | 0.73 |
