← Benchmarks
MMMLU
Multilingual Massive Multitask Language Understanding dataset released by OpenAI, featuring professionally translated MMLU test questions across 14 languages including Arabic, Bengali, German, Spanish, French, Hindi, Indonesian, Italian, Japanese, Korean, Portuguese, Swahili, Yoruba, and Chinese. Contains approximately 15,908 multiple-choice questions per language covering 57 subjects.
id mmmlu · max 1 · 51 models reported
| # | Model | Score |
|---|---|---|
| 1 | Claude Mythos Preview Anthropic | 0.93 |
| 2 | Gemini 3.1 Pro Google | 0.93 |
| 3 | Gemini 3 Flash Google | 0.92 |
| 4 | Gemini 3 Pro Google | 0.92 |
| 5 | Claude Opus 4.7 Anthropic | 0.92 |
| 6 | Claude Opus 4.6 Anthropic | 0.91 |
| 7 | Claude Opus 4.5 Anthropic | 0.91 |
| 8 | Qwen3.7 Max Alibaba Cloud / Qwen Team | 0.90 |
| 9 | GPT-5.2 OpenAI | 0.90 |
| 10 | Claude Opus 4.1 Anthropic | 0.90 |
| 11 | Qwen3.6 Plus Alibaba Cloud / Qwen Team | 0.90 |
| 12 | Claude Sonnet 4.6 Anthropic | 0.89 |
| 13 | Claude Sonnet 4.5 Anthropic | 0.89 |
| 14 | Qwen3.7-Plus Alibaba Cloud / Qwen Team | 0.89 |
| 15 | Gemini 3.1 Flash-Lite Google | 0.89 |
| 16 | Claude Opus 4 Anthropic | 0.89 |
