← Benchmarks
MT-Bench
MT-Bench is a challenging multi-turn benchmark that measures the ability of large language models to engage in coherent, informative, and engaging conversations. It uses strong LLMs as judges for scalable and explainable evaluation of multi-turn dialogue capabilities.
id mt-bench · max 100 · 13 models reported
| # | Model | Score |
|---|---|---|
| 1 | Hermes 3 70B Nous Research · open | 8.99 |
| 2 | Hermes 3 405B Nous Research · open | 8.93 |
| 3 | Qwen2.5 72B Instruct Alibaba Cloud / Qwen Team · open | 0.94 |
