← Benchmarks
SWE-Bench Multimodal
SWE-Bench Multimodal extends SWE-Bench to evaluate language models on software engineering tasks that involve visual inputs such as screenshots, UI mockups, and diagrams alongside code understanding.
id swe-bench-multimodal · max 1 · 5 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
