← Benchmarks
TAU3-Bench
TAU3-Bench is a benchmark for evaluating general-purpose agent capabilities, testing models on multi-turn interactions with simulated user models, retrieval, and complex decision-making scenarios.
id tau3-bench · max 1 · 9 models reported
| # | Model | Score |
|---|---|---|
| 1 | MiMo-V2.5-Pro Xiaomi · open | 0.73 |
| 2 | Qwen3.6 Plus Alibaba Cloud / Qwen Team | 0.71 |
| 3 | GLM-5.1 Zhipu AI · open | 0.71 |
