← Benchmarks
BrowseComp
BrowseComp is a benchmark comprising 1,266 questions that challenge AI agents to persistently navigate the internet in search of hard-to-find, entangled information. The benchmark measures agents' ability to exercise persistence in information gathering, demonstrate creativity in web navigation, and find concise, verifiable answers. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers.
id browsecomp · max 1 · 67 models reported
| # | Model | Score |
|---|---|---|
| 1 | Atria Dawn Preview Shanghai AI Laboratory · open | 0.93 |
| 2 | GPT-6 Astra OpenAI | 0.92 |
| 3 | Kimi K3 Moonshot AI · open | 0.91 |
| 4 | Claude Opus 5 Anthropic | 0.91 |
| 5 | GPT-5.6 Sol OpenAI | 0.90 |
| 6 | GPT-5.5 Pro OpenAI | 0.90 |
| 7 | GPT-5.6 Terra OpenAI | 0.88 |
| 8 | Claude Mythos Preview Anthropic | 0.87 |
| 9 | Kimi K2.6 Moonshot AI · open | 0.86 |
| 10 | Seed 2.1 Pro ByteDance | 0.86 |
| 11 | Gemini 3.1 Pro Google | 0.86 |
| 12 | Seed 2.1 Turbo ByteDance | 0.85 |
| 13 | Claude Sonnet 5 Anthropic | 0.85 |
| 14 | GPT-5.5 OpenAI | 0.84 |
| 15 | Claude Opus 4.8 Anthropic | 0.84 |
| 16 | Hy3 Tencent · open | 0.84 |
| 17 | Claude Opus 4.6 Anthropic | 0.84 |
| 18 | MiniMax M3 MiniMax · open | 0.84 |
| 19 | DeepSeek-V4-Pro-Max DeepSeek · open | 0.83 |
| 20 | GPT-5.6 Luna OpenAI | 0.83 |
| 21 | GPT-5.4 OpenAI | 0.83 |
| 22 | Claude Opus 4.7 Anthropic | 0.79 |
| 23 | GLM-5.1 Zhipu AI · open | 0.79 |
| 24 | GPT-5.2 Pro OpenAI | 0.78 |
| 25 | Inkling-Small Thinking Machines Lab · open | 0.77 |
| 26 | Seed 2.0 Pro ByteDance | 0.77 |
| 27 | MiniMax M2.5 MiniMax · open | 0.76 |
| 28 | GLM-5 Zhipu AI · open | 0.76 |
| 29 | Kimi K2.5 Moonshot AI · open | 0.75 |
| 30 | Claude Sonnet 4.6 Anthropic | 0.75 |
| 31 | DeepSeek-V4-Flash-Max DeepSeek · open | 0.73 |
| 32 | Ling 3.0 Flash InclusionAI | 0.72 |
