all the models — AI benchmark observatory
← Benchmarks

Tau-bench

τ-bench: A benchmark for tool-agent-user interaction in real-world domains. Tests language agents' ability to interact with users and follow domain-specific rules through dynamic conversations using API tools and policy guidelines across retail and airline domains. Evaluates consistency and reliability of agent behavior over multiple trials.

id tau-bench · max 1 · 6 models reported

#ModelScore
1Step-3.5-Flash
StepFun · open
0.88
2GLM-4.7
Zhipu AI · open
0.87
3MiMo-V2-Flash
Xiaomi · open
0.80
4GLM-4.7-Flash
Zhipu AI · open
0.80
5MiniMax M2
MiniMax · open
0.77