all the models — AI benchmark observatory
← Benchmarks

MT-Bench

MT-Bench is a challenging multi-turn benchmark that measures the ability of large language models to engage in coherent, informative, and engaging conversations. It uses strong LLMs as judges for scalable and explainable evaluation of multi-turn dialogue capabilities.

id mt-bench · max 100 · 13 models reported

#ModelScore
1Hermes 3 70B
Nous Research · open
8.99
2Hermes 3 405B
Nous Research · open
8.93
3Qwen2.5 72B Instruct
Alibaba Cloud / Qwen Team · open
0.94
MT-Bench Leaderboard · all the models