← Benchmarks
VLMsAreBlind
A vision-language benchmark that probes blind spots and brittle reasoning in multimodal models.
id vlmsareblind · max 1 · 6 models reported
| # | Model | Score |
|---|---|---|
| 1 | Qwen3.5-35B-A3B Alibaba Cloud / Qwen Team · open | 0.97 |
| 2 | Qwen3.6-27B Alibaba Cloud / Qwen Team · open | 0.97 |
| 3 | Qwen3.5-27B Alibaba Cloud / Qwen Team · open | 0.97 |
| 4 | Qwen3.5-122B-A10B Alibaba Cloud / Qwen Team · open | 0.97 |
| 5 | Seed 2.0 Mini ByteDance | 0.93 |
| 6 | Seed 1.8 ByteDance | 0.93 |
