all the models — AI benchmark observatory
← Benchmarks

BrowseComp

BrowseComp is a benchmark comprising 1,266 questions that challenge AI agents to persistently navigate the internet in search of hard-to-find, entangled information. The benchmark measures agents' ability to exercise persistence in information gathering, demonstrate creativity in web navigation, and find concise, verifiable answers. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers.

id browsecomp · max 1 · 67 models reported

#ModelScore
1Atria Dawn Preview
Shanghai AI Laboratory · open
0.93
2GPT-6 Astra
OpenAI
0.92
3Kimi K3
Moonshot AI · open
0.91
4Claude Opus 5
Anthropic
0.91
5GPT-5.6 Sol
OpenAI
0.90
6GPT-5.5 Pro
OpenAI
0.90
7GPT-5.6 Terra
OpenAI
0.88
8Claude Mythos Preview
Anthropic
0.87
9Kimi K2.6
Moonshot AI · open
0.86
10Seed 2.1 Pro
ByteDance
0.86
11Gemini 3.1 Pro
Google
0.86
12Seed 2.1 Turbo
ByteDance
0.85
13Claude Sonnet 5
Anthropic
0.85
14GPT-5.5
OpenAI
0.84
15Claude Opus 4.8
Anthropic
0.84
16Hy3
Tencent · open
0.84
17Claude Opus 4.6
Anthropic
0.84
18MiniMax M3
MiniMax · open
0.84
19DeepSeek-V4-Pro-Max
DeepSeek · open
0.83
20GPT-5.6 Luna
OpenAI
0.83
21GPT-5.4
OpenAI
0.83
22Claude Opus 4.7
Anthropic
0.79
23GLM-5.1
Zhipu AI · open
0.79
24GPT-5.2 Pro
OpenAI
0.78
25Inkling-Small
Thinking Machines Lab · open
0.77
26Seed 2.0 Pro
ByteDance
0.77
27MiniMax M2.5
MiniMax · open
0.76
28GLM-5
Zhipu AI · open
0.76
29Kimi K2.5
Moonshot AI · open
0.75
30Claude Sonnet 4.6
Anthropic
0.75
31DeepSeek-V4-Flash-Max
DeepSeek · open
0.73
32Ling 3.0 Flash
InclusionAI
0.72