← Benchmarks
MME
A comprehensive evaluation benchmark for Multimodal Large Language Models measuring both perception and cognition abilities across 14 subtasks. Features manually designed instruction-answer pairs to avoid data leakage and provides systematic quantitative assessment of MLLM capabilities.
id mme · max 1 · 4 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
