all the models — AI benchmark observatory
← Benchmarks

DeepSearchQA

DeepSearchQA is a benchmark for evaluating deep search and question-answering capabilities, testing models' ability to perform multi-hop reasoning and information retrieval across complex knowledge domains.

id deepsearchqa · max 1 · 11 models reported

#ModelScore
1Atria Dawn Preview
Shanghai AI Laboratory · open
0.96
2Kimi K3
Moonshot AI · open
0.95
3Claude Opus 4.8
Anthropic
0.93
4Claude Opus 4.6
Anthropic
0.91
5Hy3
Tencent · open
0.91
6Muse Spark 1.3
Meta
0.89
7MiMo-V2-Pro
Xiaomi
0.87
8Kimi K2.6
Moonshot AI · open
0.83
9Kimi K2.5
Moonshot AI · open
0.77
10Muse Spark
Meta
0.75
11Muse Glimmer-30B
Meta · open
0.75