Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

τ³-Bench Tool-Agent-User Evaluation (τ³-bench results)

τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Use each result with its attached source and setup label. Match domain or published average, task release, modality, agent and user models, scaffold, prompts, trial count, and pass^k policy before comparison. The sorted table does not imply one controlled benchmark run.

Operator receipt: 18 sourced rows are currently displayable on this page; the highest published score among these 18 models is Mercury 2.5 at 96.0%, which does not establish a market leader.

Honest limit: Ten sourced rows come from several provider reports, and their labels include telecom and broader published averages. BenchLM did not rerun them. The framework's active task fixes and expanding modalities also mean an older result may not describe the current release.

Benchmark score on τ³-bench sourced results — September 10, 2026

We mirror the published score view for τ³-bench sourced results. The public snapshot contains 18 models. Mercury 2.5 has the highest published score at 96.0%, but the available coverage and evaluation setups do not establish a market leader. We do not use these results to rank models overall.

18 modelsAgenticCurrentDisplay onlyUpdated September 10, 2026

Benchmark score table (18 models)

Score
1
Mercury 2.5Inception · Closedτ³-bench Telecom
96.0%
2
Mistral Medium 3.5 128BMistral · Open weightτ³-bench Telecom
91.4%
3
MiMo-V2.5-ProXiaomi · Closedτ³-bench published setup
72.9%
4
Nemotron 3 UltraNVIDIA · Open weightτ³-bench published average
70.9%
5
Qwen3.6 PlusAlibaba · Closedτ³-bench published setup
70.7%
6
GLM-5.1Z.AI · Open weightτ³-bench published setup
70.6%
7
Claude Opus 4.5Anthropic · Closedτ³-bench published setup
70.2%
8
Qwen3.5 397BAlibaba · Open weightτ³-bench published setup
68.4%
9
Qwen3.6-35B-A3BAlibaba · Open weightτ³-bench published setup
67.2%
10
Pokee-Isaac 28BPokee AI · Closedτ³-bench Telecom
66.2%
11
Kimi K2.5Moonshot AI · Open weightτ³-bench published setup
65.7%
12
GLM-5Z.AI · Open weightτ³-bench published setup
65.6%
13
Granite 4.2 30BIBM · Open weightτ³-bench published setup
62.0%
14
Granite 4.2 8BIBM · Open weightτ³-bench published average
58.1%
15
Granite 4.2 3BIBM · Open weightτ³-bench published setup
45.8%
16
Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA · Open weightτ³-bench published setup
9.5%
17
Mercury 2Inception · Closedτ³-bench published setup
9.0%
18
LFM2.5-2.6BLiquidAI · Open weightτ³-bench published setup
5.7%

The published τ³-bench results snapshot places Mercury 2.5 first at 96.0%. The third row is 23.1 points behind. The broader top-10 range is 29.8 points, so the table still separates the published systems.

18 models have been evaluated on τ³-bench results. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. τ³-bench results is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About τ³-bench results

Year

2026

Tasks

Corrected customer-service tasks plus knowledge and voice evaluation modes

Format

Published domain or average success results

Difficulty

Long-horizon, multimodal, and knowledge-aware tool use

The maintained repository now identifies the framework as τ³-bench and documents text, voice, telecom, airline, retail, and knowledge-aware banking evaluation. BenchLM keeps provider-published τ³ rows separate from original TAU-bench and τ²-bench rows.

BenchLM freshness & provenance

Version

τ³-bench results 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

Is τ³-bench the same as original TAU-bench?

No. It is the maintained successor framework with corrected tasks and newer domains and modalities. BenchLM keeps original TAU-bench, τ²-bench, τ² Airline, and τ³-bench on separate routes.

Can every τ³-bench result be ranked together?

No. Match the domain or average, release, modality, agent and user models, scaffold, prompts, trials, and pass^k policy. Provider tables without those same controls are useful source receipts, not one apples-to-apples leaderboard.

Last updated: September 10, 2026 · BenchLM version τ³-bench results 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.