← Benchmarks
Claw-Eval
Claw-Eval tests real-world agentic task completion across complex multi-step scenarios, evaluating a model's ability to use tools, navigate environments, and complete end-to-end tasks autonomously.
id claw-eval · max 1 · 14 models reported
| # | Model | Score |
|---|---|---|
| 1 | Kimi K2.6 Moonshot AI · open | 0.81 |
| 2 | GLM-5V-Turbo Zhipu AI | 0.75 |
| 3 | MiniMax M3 MiniMax · open | 0.74 |
