all the models — AI benchmark observatory
← Benchmarks

MMBench

A bilingual benchmark for assessing multi-modal capabilities of vision-language models through multiple-choice questions in both English and Chinese, providing systematic evaluation across diverse vision-language tasks with robust metrics.

id mmbench · max 1 · 9 models reported

#ModelScore
1Step3-VL-10B
StepFun · open
0.92
2Qwen2.5 VL 72B Instruct
Alibaba Cloud / Qwen Team · open
0.88
3Phi-4-multimodal-instruct
Microsoft · open
0.87
4Qwen2-VL-72B-Instruct
Alibaba Cloud / Qwen Team · open
0.86
MMBench Leaderboard · all the models