← Benchmarks
VideoMME w sub.
The first-ever comprehensive evaluation benchmark of Multi-modal LLMs in Video analysis. Features 900 videos (254 hours) with 2,700 question-answer pairs covering 6 primary visual domains and 30 subfields. Evaluates temporal understanding across short (11 seconds) to long (1 hour) videos with multi-modal inputs including video frames, subtitles, and audio.
id videomme-w-sub. · max 1 · 12 models reported
| # | Model | Score |
|---|---|---|
| 1 | Qwen3.8 Max Alibaba Cloud / Qwen Team · open | 0.90 |
| 2 | Seed 1.8 ByteDance | 0.88 |
| 3 | Qwen3.6-27B Alibaba Cloud / Qwen Team · open | 0.88 |
| 4 | Qwen3.5-122B-A10B Alibaba Cloud / Qwen Team · open | 0.87 |
| 5 | Qwen3.5-27B Alibaba Cloud / Qwen Team · open | 0.87 |
| 6 | GPT-5 OpenAI | 0.87 |
| 7 | Qwen3.5-35B-A3B Alibaba Cloud / Qwen Team · open | 0.87 |
| 8 | Qwen3.6-35B-A3B Alibaba Cloud / Qwen Team · open | 0.87 |
