all the models — AI benchmark observatory
← Benchmarks

Tau2 Retail

τ²-bench retail domain evaluates conversational AI agents in customer service scenarios within a dual-control environment where both agent and user can interact with tools. Tests tool-agent-user interaction, rule adherence, and task consistency in retail customer support contexts.

id tau2-retail · max 1 · 27 models reported

#ModelScore

No scores for this benchmark yet.

Tau2 Retail Leaderboard · all the models