Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Tool-Agent-User Benchmark (TAU-bench)

Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.

Data verified 25 confirmed releases in the last 30 daysStart free brief

How to read this benchmark

Editorial review by Glevd · 2026-07-15

Read this route as the owner of original 2024 TAU-bench methodology. Compare a result only when the domain, archived task release, agent and user models, prompting strategy, tools, number of trials, and pass^k metric are identified. Do not merge it with τ² or τ³ results.

Operator receipt: 0 sourced rows are currently displayable on this page.

Honest limit: BenchLM stores 38 raw numbers under the original key, but none has an exact source attachment or the domain and pass^k labels needed for a valid comparison. The route therefore publishes no score table. The archived repository also warns that its airline and retail tasks are outdated.

About TAU-bench

Year

2024

Tasks

Airline and retail task sets in the archived 2024 release

Format

Domain-specific pass^1 through pass^4 task success

Difficulty

Policy-constrained, multi-turn customer service

The original release reports airline and retail separately and measures reliability with pass^1 through pass^4 across repeated trials. Its repository now warns that those task files are outdated and directs evaluators to the maintained τ³-bench repository.

BenchLM freshness & provenance

Version

TAU-bench 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

Why is there no original TAU-bench score table?

The 38 stored raw numbers do not have exact source attachments or the domain and pass^k labels required for a valid comparison. BenchLM withholds them until those receipts are attached instead of presenting unsupported values as a leaderboard.

Is original TAU-bench still the current release?

No. The original repository warns that its airline and retail task files are outdated and points evaluators to the maintained τ³-bench repository for corrected tasks and newer evaluation modes.

What does original TAU-bench measure?

It measures whether an agent can complete airline or retail customer-service tasks through multi-turn conversation and domain tools while following policy. The paper reports domain-specific pass^k reliability across repeated trials.

Last updated: August 21, 2026 · BenchLM version TAU-bench 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.