all the models — AI benchmark observatory
← Benchmarks

MM-MT-Bench

A multi-turn LLM-as-a-judge evaluation benchmark for testing multimodal instruction-tuned models' ability to follow user instructions in multi-turn dialogues and answer open-ended questions in a zero-shot manner.

id mm-mt-bench · max 100 · 17 models reported

#ModelScore

No scores for this benchmark yet.