TAU-bench
AI benchmark
Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.
- Domains
- Agentic
- Data source
- benchlm-benchmarks
- Catalogued
Catalogue entry from the RL Engineering daily scrape of public sources. Data reflects the 2026-08-30 snapshot.