← Benchmarks
SWE-MM
SWE-MM evaluates software-engineering agents on repository tasks that require understanding both source code and visual evidence.
id swe-mm · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
SWE-MM evaluates software-engineering agents on repository tasks that require understanding both source code and visual evidence.
id swe-mm · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.