all the models — AI benchmark observatory
← Benchmarks

SWE-Marathon

SWE-Marathon is an ultra-long-horizon software engineering benchmark covering tasks such as building compilers, optimizing kernels, and developing production-grade services. It measures whether agents can sustain quality across extremely long engineering trajectories.

id swe-marathon · max 1 · 6 models reported

#ModelScore

No scores for this benchmark yet.