← Benchmarks
MMBench
A bilingual benchmark for assessing multi-modal capabilities of vision-language models through multiple-choice questions in both English and Chinese, providing systematic evaluation across diverse vision-language tasks with robust metrics.
id mmbench · max 1 · 9 models reported
| # | Model | Score |
|---|---|---|
| 1 | Step3-VL-10B StepFun · open | 0.92 |
| 2 | Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team · open | 0.88 |
| 3 | Phi-4-multimodal-instruct Microsoft · open | 0.87 |
| 4 | Qwen2-VL-72B-Instruct Alibaba Cloud / Qwen Team · open | 0.86 |
