← Benchmarks
Video-MME
Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos totaling 254 hours with 2,700 human-annotated question-answer pairs across 6 primary visual domains (Knowledge, Film & Television, Sports Competition, Life Record, Multilingual, and others) and 30 subfields. The benchmark evaluates models across diverse temporal dimensions (11 seconds to 1 hour), integrates multi-modal inputs including video frames, subtitles, and audio, and uses rigorous manual labeling by expert annotators for precise assessment.
id video-mme · max 1 · 17 models reported
| # | Model | Score |
|---|---|---|
| 1 | Seed 2.1 Pro ByteDance | 0.89 |
| 2 | Seed 2.1 Turbo ByteDance | 0.89 |
| 3 | Qwen3.7-Plus Alibaba Cloud / Qwen Team | 0.88 |
| 4 | MiMo-V2.5 Xiaomi · open | 0.88 |
| 5 | Kimi K2.5 Moonshot AI · open | 0.87 |
| 6 | MiniMax M3 MiniMax · open | 0.85 |
