all the models — AI benchmark observatory
← Benchmarks

Arena Hard

Arena-Hard-Auto is an automatic evaluation benchmark for instruction-tuned LLMs consisting of 500 challenging real-world prompts curated by BenchBuilder. It includes open-ended software engineering problems, mathematical questions, and creative writing tasks. The benchmark uses LLM-as-a-Judge methodology with GPT-4.1 and Gemini-2.5 as automatic judges to approximate human preference. Arena-Hard achieves 98.6% correlation with human preference rankings and provides 3x higher separation of model performances compared to MT-Bench, making it highly effective for distinguishing between models of similar quality.

id arena-hard · max 1 · 27 models reported

#ModelScore
1Qwen3 235B A22B
Alibaba Cloud / Qwen Team · open
0.96
2Qwen3 32B
Alibaba Cloud / Qwen Team · open
0.94
3Qwen3 14B
Alibaba Cloud / Qwen Team · open
0.92
4Qwen3 30B A3B
Alibaba Cloud / Qwen Team · open
0.91
5Llama-3.3 Nemotron Super 49B v1
NVIDIA · open
0.88
6Mistral Small 3 24B Instruct
Mistral AI · open
0.88
7Qwen2.5 72B Instruct
Alibaba Cloud / Qwen Team · open
0.81
8Phi 4 Reasoning Plus
Microsoft · open
0.79
9DeepSeek-V2.5
DeepSeek · open
0.76
10Phi 4
Microsoft · open
0.75
11Phi 4 Reasoning
Microsoft · open
0.73
12Ministral 8B Instruct
Mistral AI · open
0.71
13Jamba 1.5 Large
AI21 Labs · open
0.65
14Mistral Small 4
Mistral AI · open
0.58
15Granite 3.3 8B Base
IBM · open
0.58
16Granite 3.3 8B Instruct
IBM · open
0.58
17MiniStral 3 (14B Instruct 2512)
Mistral AI · open
0.55
18Mistral Large 3
Mistral AI · open
0.55
19Qwen2.5 7B Instruct
Alibaba Cloud / Qwen Team · open
0.52
20Ministral 3 (8B Instruct 2512)
Mistral AI · open
0.51
21Jamba 1.5 Mini
AI21 Labs · open
0.46
22Mistral Small 3.2 24B Instruct
Mistral AI · open
0.43
23Phi-3.5-MoE-instruct
Microsoft · open
0.38
24Phi-3.5-mini-instruct
Microsoft · open
0.37
25Phi 4 Mini
Microsoft · open
0.33
26Ministral 3 (3B Instruct 2512)
Mistral AI · open
0.30
27IBM Granite 4.0 Tiny Preview
IBM · open
0.27