all the models — AI benchmark observatory
← Benchmarks

AGIEval

A human-centric benchmark for evaluating foundation models on standardized exams including college entrance exams (Gaokao, SAT), law school admission tests (LSAT), math competitions, lawyer qualification tests, and civil service exams. Contains 20 tasks (18 multiple-choice, 2 cloze) designed to assess understanding, knowledge, reasoning, and calculation abilities in real-world academic and professional contexts.

id agieval · max 1 · 11 models reported

#ModelScore

No scores for this benchmark yet.

AGIEval Leaderboard · all the models