all the models — AI benchmark observatory
← Benchmarks

TAU-bench Airline

Part of τ-bench (TAU-bench), a benchmark for Tool-Agent-User interaction in real-world domains. The airline domain evaluates language agents' ability to interact with users through dynamic conversations while following domain-specific rules and using API tools. Agents must handle airline-related tasks and policies reliably.

id tau-bench-airline · max 1 · 23 models reported

#ModelScore

No scores for this benchmark yet.

TAU-bench Airline Leaderboard · all the models