all the models — AI benchmark observatory
← Benchmarks

VLMsAreBlind

A vision-language benchmark that probes blind spots and brittle reasoning in multimodal models.

id vlmsareblind · max 1 · 6 models reported

#ModelScore
1Qwen3.5-35B-A3B
Alibaba Cloud / Qwen Team · open
0.97
2Qwen3.6-27B
Alibaba Cloud / Qwen Team · open
0.97
3Qwen3.5-27B
Alibaba Cloud / Qwen Team · open
0.97
4Qwen3.5-122B-A10B
Alibaba Cloud / Qwen Team · open
0.97
5Seed 2.0 Mini
ByteDance
0.93
6Seed 1.8
ByteDance
0.93