← Benchmarks
Tau-bench
τ-bench: A benchmark for tool-agent-user interaction in real-world domains. Tests language agents' ability to interact with users and follow domain-specific rules through dynamic conversations using API tools and policy guidelines across retail and airline domains. Evaluates consistency and reliability of agent behavior over multiple trials.
id tau-bench · max 1 · 6 models reported
| # | Model | Score |
|---|---|---|
| 1 | Step-3.5-Flash StepFun · open | 0.88 |
| 2 | GLM-4.7 Zhipu AI · open | 0.87 |
| 3 | MiMo-V2-Flash Xiaomi · open | 0.80 |
| 4 | GLM-4.7-Flash Zhipu AI · open | 0.80 |
| 5 | MiniMax M2 MiniMax · open | 0.77 |
