← Benchmarks
WideSearch
WideSearch is an agentic search benchmark that evaluates models' ability to perform broad, parallel search operations across multiple sources. It tests wide-coverage information retrieval and synthesis capabilities.
id widesearch · max 1 · 15 models reported
| # | Model | Score |
|---|---|---|
| 1 | Hy4 preview Tencent · open | 0.84 |
| 2 | Atria Dawn Preview Shanghai AI Laboratory · open | 0.82 |
| 3 | Qwen3.8 Max Alibaba Cloud / Qwen Team · open | 0.82 |
| 4 | Kimi K2.6 Moonshot AI · open | 0.81 |
| 5 | Kimi K2.5 Moonshot AI · open | 0.79 |
| 6 | Hy3 Tencent · open | 0.76 |
| 7 | Qwen3.6 Plus Alibaba Cloud / Qwen Team | 0.74 |
| 8 | Qwen3.5-397B-A17B Alibaba Cloud / Qwen Team · open | 0.74 |
| 9 | Ling 3.0 Flash InclusionAI | 0.74 |
