all the models — AI benchmark observatory
← Benchmarks

Toolathlon-Verified

Verified Toolathlon evaluation of agent tool-use capability, as reported in the MiMo-V2.6 release.

id toolathlon-verified · max 1 · 3 models reported

#ModelScore
1Claude Opus 5.5
Anthropic
0.78
2MiMo-V2.6-Pro
Xiaomi · open
0.77
3MiMo-V2.6-Flash
Xiaomi · open
0.74
Toolathlon-Verified Leaderboard · all the models