all the models — AI benchmark observatory
← Benchmarks

MME

A comprehensive evaluation benchmark for Multimodal Large Language Models measuring both perception and cognition abilities across 14 subtasks. Features manually designed instruction-answer pairs to avoid data leakage and provides systematic quantitative assessment of MLLM capabilities.

id mme · max 1 · 4 models reported

#ModelScore

No scores for this benchmark yet.

MME Leaderboard · all the models