all the models — AI benchmark observatory
← Benchmarks

TAU-bench Retail

A benchmark for evaluating tool-agent-user interaction in retail environments. Tests language agents' ability to handle dynamic conversations with users while using domain-specific API tools and following policy guidelines. Evaluates agents on tasks like order cancellations, address changes, and order status checks through multi-turn conversations.

id tau-bench-retail · max 1 · 25 models reported

#ModelScore

No scores for this benchmark yet.

TAU-bench Retail Leaderboard · all the models