# τ³-Bench Tool-Agent-User Evaluation (τ³-bench results)

> τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.

Canonical page: https://benchlm.ai/benchmarks/tau3-bench

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About τ³-bench results

- Year: 2026
- Tasks: Corrected customer-service tasks plus knowledge and voice evaluation modes
- Format: Published domain or average success results
- Difficulty: Long-horizon, multimodal, and knowledge-aware tool use
- Paper: [Official τ³-bench repository and release notes](https://github.com/sierra-research/tau2-bench)

The maintained repository now identifies the framework as τ³-bench and documents text, voice, telecom, airline, retail, and knowledge-aware banking evaluation. BenchLM keeps provider-published τ³ rows separate from original TAU-bench and τ²-bench rows.

τ³-bench results is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (17 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Mercury 2.5](/models/mercury-2-5) | Inception | 96.0% |
| 2 | [Mistral Medium 3.5 128B](/models/mistral-medium-3-5-128b) | Mistral | 91.4% |
| 3 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | Xiaomi | 72.9% |
| 4 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 70.9% |
| 5 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 70.7% |
| 6 | [GLM-5.1](/models/glm-5-1) | Z.AI | 70.6% |
| 7 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 70.2% |
| 8 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 68.4% |
| 9 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 67.2% |
| 10 | [Pokee-Isaac 28B](/models/pokee-isaac-28b) | Pokee AI | 66.2% |
| 11 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 65.7% |
| 12 | [GLM-5](/models/glm-5) | Z.AI | 65.6% |
| 13 | [Granite 4.2 30B](/models/granite-4-2-30b) | IBM | 62.0% |
| 14 | [Granite 4.2 8B](/models/granite-4-2-8b) | IBM | 58.1% |
| 15 | [Granite 4.2 3B](/models/granite-4-2-3b) | IBM | 45.8% |
| 16 | [Nemotron 3.5 Lightning 30B A3B NVFP4](/models/nemotron-3-5-lightning-30b-a3b-nvfp4) | NVIDIA | 9.5% |
| 17 | [LFM2.5-2.6B](/models/lfm2-5-2-6b) | LiquidAI | 5.7% |

## FAQ

### Is τ³-bench the same as original TAU-bench?

No. It is the maintained successor framework with corrected tasks and newer domains and modalities. BenchLM keeps original TAU-bench, τ²-bench, τ² Airline, and τ³-bench on separate routes.

### Can every τ³-bench result be ranked together?

No. Match the domain or average, release, modality, agent and user models, scaffold, prompts, trials, and pass^k policy. Provider tables without those same controls are useful source receipts, not one apples-to-apples leaderboard.

## Compare Top Models on τ³-bench results

- [Mercury 2.5 vs Mistral Medium 3.5 128B](/compare/mercury-2-5-vs-mistral-medium-3-5-128b)
- [Mistral Medium 3.5 128B vs MiMo-V2.5-Pro](/compare/mimo-v2-5-pro-vs-mistral-medium-3-5-128b)
- [MiMo-V2.5-Pro vs Nemotron 3 Ultra](/compare/mimo-v2-5-pro-vs-nemotron-3-ultra)
- [Nemotron 3 Ultra vs Qwen3.6 Plus](/compare/nemotron-3-ultra-vs-qwen3-6-plus)
