Tool-Agent-User Benchmark (TAU-bench)
Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.
Data verified 31 confirmed releases in the last 30 daysSee provider release alertsHow to read this benchmark
Editorial review by Glevd · 2026-07-15
Read this route as the owner of original 2024 TAU-bench methodology. Compare a result only when the domain, archived task release, agent and user models, prompting strategy, tools, number of trials, and pass^k metric are identified. Do not merge it with τ² or τ³ results.
Operator receipt: 0 sourced rows are currently displayable on this page.
Honest limit: BenchLM stores 38 raw numbers under the original key, but none has an exact source attachment or the domain and pass^k labels needed for a valid comparison. The route therefore publishes no score table. The archived repository also warns that its airline and retail tasks are outdated.
About TAU-bench
Year
2024
Tasks
Airline and retail task sets in the archived 2024 release
Format
Domain-specific pass^1 through pass^4 task success
Difficulty
Policy-constrained, multi-turn customer service
The original release reports airline and retail separately and measures reliability with pass^1 through pass^4 across repeated trials. Its repository now warns that those task files are outdated and directs evaluators to the maintained τ³-bench repository.
BenchLM freshness & provenance
Version
TAU-bench 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
Why is there no original TAU-bench score table?
The 38 stored raw numbers do not have exact source attachments or the domain and pass^k labels required for a valid comparison. BenchLM withholds them until those receipts are attached instead of presenting unsupported values as a leaderboard.
Is original TAU-bench still the current release?
No. The original repository warns that its airline and retail task files are outdated and points evaluators to the maintained τ³-bench repository for corrected tasks and newer evaluation modes.
What does original TAU-bench measure?
It measures whether an agent can complete airline or retail customer-service tasks through multi-turn conversation and domain tools while following policy. The paper reports domain-specific pass^k reliability across repeated trials.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.