τ²-Bench Tool-Agent-User Evaluation (τ²-bench results)
We show this table for reference; we do not rank on it.
This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.
Benchmark score on τ²-bench sourced results — October 2, 2026
We compile the τ²-bench sourced results rows from provider self-reports and secondary reports. The table contains 152 models. GLM-5.2 has the highest published score at 99.1%, but the available coverage and evaluation setups do not establish a market leader. We do not use these results to rank models overall.
GLM-5.2
Z.AI
Artificial Analysis τ²-bench
GPT-5.4
OpenAI
τ²-bench Telecom
GLM-4.7-Flash
Z.AI
Artificial Analysis τ²-bench
152 modelsAgenticCurrentDisplay onlyUpdated October 2, 2026
Benchmark score table (152 models)
ScoreHow to read this leaderboard
Editorial review by Glevd · 2026-07-15
Use a row only with its attached source and setup label. Match the domain, task release, agent model, user-simulator model, scaffold, prompts, trial count, and pass^k metric before comparing scores. The sorted table is a source ledger, not a controlled cross-provider ranking.
Operator receipt: 152 sourced rows are currently displayable on this page; the highest published score among these 152 models is GLM-5.2 at 99.1%, which does not establish a market leader.
Honest limit: The page mixes a large third-party telecom snapshot with smaller provider-published slices. Those sources do not use one guaranteed-common harness or reporting policy, and BenchLM did not rerun them. A higher number can reflect a different user model, prompt, domain, task release, or repeat policy.
Among the reported τ²-bench results rows, GLM-5.2 is first at 99.1%. The third row is 0.3 points behind. The broader top-10 range is 1.4 points, so many of the published results sit in a relatively narrow band.
152 models have been evaluated on τ²-bench results. The benchmark falls in the Agentic category. τ²-bench results is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About τ²-bench results
Year
2025
Tasks
Airline, retail, and telecom customer-service task sets
Format
Published domain success or pass^k results
Difficulty
Dual-control customer-service workflows
τ²-bench extends the original benchmark with a dual-control telecom domain where the agent and simulated user can both act through tools. The maintained framework also includes airline and retail. BenchLM keeps each exact source attached and labels the published setup instead of treating every row as one controlled run.
Freshness and provenance
Version
τ²-Bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
Are all τ²-bench scores directly comparable?
No. Match the domain, task release, agent model, user model, scaffold, prompts, number of trials, and pass^k definition. BenchLM labels each sourced row so a telecom result or third-party implementation is not silently treated as the same setup as an airline, retail, or aggregate result.
Does a high τ²-bench score prove production support reliability?
No. It measures success in simulated customer-service environments under a reported setup. Production identity checks, permission boundaries, changing policies, latency, cost, monitoring, and human escalation still need separate testing.
Compare top models on τ²-bench results
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.