all the models — AI benchmark observatory
← Benchmarks

Claw-Eval

Claw-Eval tests real-world agentic task completion across complex multi-step scenarios, evaluating a model's ability to use tools, navigate environments, and complete end-to-end tasks autonomously.

id claw-eval · max 1 · 14 models reported

#ModelScore
1Kimi K2.6
Moonshot AI · open
0.81
2GLM-5V-Turbo
Zhipu AI
0.75
3MiniMax M3
MiniMax · open
0.74
Claw-Eval Leaderboard · all the models