all the models — AI benchmark observatory
← Benchmarks

MMLU-Base

Base version of the Massive Multitask Language Understanding benchmark, evaluating language models across 57 tasks including elementary mathematics, US history, computer science, law, and other professional and academic subjects. Designed to comprehensively measure the breadth and depth of a model's academic and professional understanding.

id mmlu-base · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

MMLU-Base Leaderboard · all the models