← Benchmarks
DeepSearchQA
DeepSearchQA is a benchmark for evaluating deep search and question-answering capabilities, testing models' ability to perform multi-hop reasoning and information retrieval across complex knowledge domains.
id deepsearchqa · max 1 · 11 models reported
| # | Model | Score |
|---|---|---|
| 1 | Atria Dawn Preview Shanghai AI Laboratory · open | 0.96 |
| 2 | Kimi K3 Moonshot AI · open | 0.95 |
| 3 | Claude Opus 4.8 Anthropic | 0.93 |
| 4 | Claude Opus 4.6 Anthropic | 0.91 |
| 5 | Hy3 Tencent · open | 0.91 |
| 6 | Muse Spark 1.3 Meta | 0.89 |
| 7 | MiMo-V2-Pro Xiaomi | 0.87 |
| 8 | Kimi K2.6 Moonshot AI · open | 0.83 |
| 9 | Kimi K2.5 Moonshot AI · open | 0.77 |
| 10 | Muse Spark Meta | 0.75 |
| 11 | Muse Glimmer-30B Meta · open | 0.75 |
