all the models — AI benchmark observatory
← Benchmarks

MMLU Chat

Chat-format variant of the Massive Multitask Language Understanding benchmark, evaluating language models across 57 tasks including elementary mathematics, US history, computer science, law, and other professional and academic subjects. This version uses conversational prompting format for model evaluation.

id mmlu-chat · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

MMLU Chat Leaderboard · all the models