all the models — AI benchmark observatory
← Benchmarks

SWE-MM

SWE-MM evaluates software-engineering agents on repository tasks that require understanding both source code and visual evidence.

id swe-mm · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

SWE-MM Leaderboard · all the models