all the models — AI benchmark observatory
← Benchmarks

SWE-Bench Multimodal

SWE-Bench Multimodal extends SWE-Bench to evaluate language models on software engineering tasks that involve visual inputs such as screenshots, UI mockups, and diagrams alongside code understanding.

id swe-bench-multimodal · max 1 · 5 models reported

#ModelScore

No scores for this benchmark yet.