# Tool-Agent-User Benchmark (TAU-bench)

> Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.

Canonical page: https://benchlm.ai/benchmarks/tau-bench

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About TAU-bench

- Year: 2024
- Tasks: Airline and retail task sets in the archived 2024 release
- Format: Domain-specific pass^1 through pass^4 task success
- Difficulty: Policy-constrained, multi-turn customer service
- Paper: [τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains](https://arxiv.org/abs/2406.12045)

The original release reports airline and retail separately and measures reliability with pass^1 through pass^4 across repeated trials. Its repository now warns that those task files are outdated and directs evaluators to the maintained τ³-bench repository.

TAU-bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### Why is there no original TAU-bench score table?

The 38 stored raw numbers do not have exact source attachments or the domain and pass^k labels required for a valid comparison. BenchLM withholds them until those receipts are attached instead of presenting unsupported values as a leaderboard.

### Is original TAU-bench still the current release?

No. The original repository warns that its airline and retail task files are outdated and points evaluators to the maintained τ³-bench repository for corrected tasks and newer evaluation modes.

### What does original TAU-bench measure?

It measures whether an agent can complete airline or retail customer-service tasks through multi-turn conversation and domain tools while following policy. The paper reports domain-specific pass^k reliability across repeated trials.
