← Benchmarks
MVBench
A comprehensive multi-modal video understanding benchmark covering 20 challenging video tasks that require temporal understanding beyond single-frame analysis. Tasks span from perception to cognition, including action recognition, temporal reasoning, spatial reasoning, object interaction, scene transition, and counterfactual inference. Uses a novel static-to-dynamic method to systematically generate video tasks from existing annotations.
id mvbench · max 1 · 18 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
