← Benchmarks
MM-MT-Bench
A multi-turn LLM-as-a-judge evaluation benchmark for testing multimodal instruction-tuned models' ability to follow user instructions in multi-turn dialogues and answer open-ended questions in a zero-shot manner.
id mm-mt-bench · max 100 · 17 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
