all the models — AI benchmark observatory
← Benchmarks

C-Eval

C-Eval is a comprehensive Chinese evaluation suite designed to assess advanced knowledge and reasoning abilities of foundation models in a Chinese context. It comprises 13,948 multiple-choice questions across 52 diverse disciplines spanning humanities, science, and engineering, with four difficulty levels: middle school, high school, college, and professional. The benchmark includes C-Eval Hard, a subset of very challenging subjects requiring advanced reasoning abilities.

id c-eval · max 1 · 20 models reported

#ModelScore
1Qwen3 Max Thinking
Alibaba Cloud / Qwen Team
0.94
2Qwen3.6 Plus
Alibaba Cloud / Qwen Team
0.93
3Qwen3.5-397B-A17B
Alibaba Cloud / Qwen Team · open
0.93
4Kimi K2 Base
Moonshot AI · open
0.93