← Benchmarks
Video-MME (long, no subtitles)
Video-MME is the first-ever comprehensive evaluation benchmark for Multi-modal Large Language Models (MLLMs) in video analysis. This variant focuses on long-term videos (30min-60min) without subtitle inputs, testing robust contextual dynamics across 6 primary visual domains with 30 subfields including knowledge, film & television, sports competition, life record, and multilingual content.
id video-mme-(long,-no-subtitles) · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
