Skip to main content
BenchLM
Data

τ²-Bench Tool-Agent-User Evaluation (τ²-bench results)

We show this table for reference; we do not rank on it.

Data verified 38 confirmed releases in the last 30 daysFollow model changes

This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.

Benchmark score on τ²-bench sourced results — October 2, 2026

We compile the τ²-bench sourced results rows from provider self-reports and secondary reports. The table contains 152 models. GLM-5.2 has the highest published score at 99.1%, but the available coverage and evaluation setups do not establish a market leader. We do not use these results to rank models overall.

152 modelsAgenticCurrentDisplay onlyUpdated October 2, 2026

Benchmark score table (152 models)

Score
1
GLM-5.2Z.AI · Open weightArtificial Analysis τ²-bench
99.1%
2
GPT-5.4OpenAI · Closedτ²-bench Telecom
98.9%
3
GLM-4.7-FlashZ.AI · Open weightArtificial Analysis τ²-bench
98.8%
4
Claude Fable 5Anthropic · ClosedArtificial Analysis τ²-bench
98.5%
5
GLM-5-TurboZ.AI · ClosedArtificial Analysis τ²-bench
98.5%
6
GLM-5V-TurboZ.AI · ClosedArtificial Analysis τ²-bench
98.5%
7
Step 3.7 FlashStepFun · Open weightArtificial Analysis τ²-bench
98.5%
8
GLM-5Z.AI · Open weightArtificial Analysis τ²-bench
98.2%
9
GPT-5.5OpenAI · Closedτ²-bench Telecom
98%
10
GLM-5.1Z.AI · Open weightArtificial Analysis τ²-bench
97.7%
11
Qwen3.6 PlusAlibaba · ClosedArtificial Analysis τ²-bench
97.7%
12
Grok 4.3xAI · ClosedArtificial Analysis τ²-bench
97.7%
13
DeepSeek V4 Pro 0813DeepSeek · Open weightArtificial Analysis τ²-bench
96.2%
14
Kimi K2.6Moonshot AI · Open weightArtificial Analysis τ²-bench
95.9%
15
Kimi K2.5Moonshot AI · Open weightArtificial Analysis τ²-bench
95.9%
16
GLM-4.7Z.AI · Open weightArtificial Analysis τ²-bench
95.9%
17
Qwen 3.6 Max (preview)Alibaba · ClosedArtificial Analysis τ²-bench
95.9%
18
Kimi K2.5 (Reasoning)Moonshot AI · ClosedArtificial Analysis τ²-bench
95.9%
19
Gemini 3.1 ProGoogle · Closedτ²-bench published setup
95.6%
20
Gemini 3.5 FlashGoogle · ClosedArtificial Analysis τ²-bench
95.3%
21
Qwen3.6-35B-A3BAlibaba · Open weightArtificial Analysis τ²-bench
95.3%
22
MiniMax M2.5MiniMax · ClosedArtificial Analysis τ²-bench
95.3%
23
MiMo-V2-ProXiaomi · ClosedArtificial Analysis τ²-bench
95%
24
Qwen3.7 MaxAlibaba · ClosedArtificial Analysis τ²-bench
94.7%
25
Claude Opus 4.8Anthropic · ClosedArtificial Analysis τ²-bench
94.4%
26
Mistral Medium 3.5 128BMistral · Open weightArtificial Analysis τ²-bench
94.2%
27
Qwen3.6-27BAlibaba · Open weightArtificial Analysis τ²-bench
94.2%
28
MiMo-V2.5-ProXiaomi · ClosedArtificial Analysis τ²-bench
94.2%
29
DeepSeek V4 Pro (High)DeepSeek · Open weightArtificial Analysis τ²-bench
94.2%
30
Qwen3.5-27BAlibaba · Open weightArtificial Analysis τ²-bench
93.9%
31
Qwen3.5-122B-A10BAlibaba · Open weightArtificial Analysis τ²-bench
93.6%
32
GPT-5.4 miniOpenAI · Closedτ²-bench Telecom
93.4%
33
Grok 4.1 Fast (Reasoning)xAI · ClosedArtificial Analysis τ²-bench
93.3%
34
Qwen3.7 PlusAlibaba · ClosedArtificial Analysis τ²-bench
93%
35
GPT-5.4 nanoOpenAI · Closedτ²-bench Telecom
92.5%
36
GPT-5.2-CodexOpenAI · ClosedArtificial Analysis τ²-bench
92.1%
37
Claude Opus 4.6 (Adaptive)Anthropic · ClosedArtificial Analysis τ²-bench
92.1%
38
Muse SparkMeta · Closedτ²-bench published setup
91.5%
39
MiMo-V2-OmniXiaomi · ClosedArtificial Analysis τ²-bench
91.2%
40
Kimi K2.7 CodeMoonshot AI · Open weightArtificial Analysis τ²-bench
90.1%
41
Trinity-Large-ThinkingArcee AI · Open weightArtificial Analysis τ²-bench
90.1%
42
Trinity-Large-PreviewArcee AI · Open weightArtificial Analysis τ²-bench
90.1%
43
Claude Opus 4.5 ThinkingAnthropic · ClosedArtificial Analysis τ²-bench
89.5%
44
Qwen3.5-35B-A3BAlibaba · Open weightArtificial Analysis τ²-bench
89.2%
45
MiniMax M3MiniMax · Open weightArtificial Analysis τ²-bench
88.9%
46
Claude Opus 4.7 (Adaptive)Anthropic · ClosedArtificial Analysis τ²-bench
88.6%
47
LFM2.5-8B-A1BLiquidAI · Open weightτ²-bench Telecom
88.1%
48
Step 3.5 FlashStepFun · Open weightArtificial Analysis τ²-bench
87.4%
49
Gemini 3 ProGoogle · ClosedArtificial Analysis τ²-bench
87.1%
50
GPT-5 (medium)OpenAI · ClosedArtificial Analysis τ²-bench
86.5%
51
GPT-5.6 TerraOpenAI · ClosedArtificial Analysis τ²-bench
86.3%
52
Claude Opus 4.5Anthropic · ClosedArtificial Analysis τ²-bench
86.3%
53
Solar Pro 3Upstage · ClosedArtificial Analysis τ²-bench
86.3%
54
GPT-5.3 CodexOpenAI · ClosedArtificial Analysis τ²-bench
86%
55
Ling 2.6 FlashInclusionAI · Open weightArtificial Analysis τ²-bench
86%
56
GPT-5.3-Codex-SparkOpenAI · ClosedArtificial Analysis τ²-bench
86%
57
GPT-5.6 SolOpenAI · ClosedArtificial Analysis τ²-bench
85.1%
58
Command A+Cohere · Open weightτ²-bench Telecom
85%
59
Claude Opus 4.6Anthropic · ClosedArtificial Analysis τ²-bench
84.8%
60
MiniMax M2.7MiniMax · Open weightArtificial Analysis τ²-bench
84.8%
61
GPT-5.2OpenAI · ClosedArtificial Analysis τ²-bench
84.8%
62
GPT-5 (high)OpenAI · ClosedArtificial Analysis τ²-bench
84.8%
63
Qwen3.5 397BAlibaba · Open weightArtificial Analysis τ²-bench
83.9%
64
MiMo-V2-FlashXiaomi · Open weightArtificial Analysis τ²-bench
83.9%
65
Qwen3.5 397B (Reasoning)Alibaba · Open weightArtificial Analysis τ²-bench
83.9%
66
Nemotron 3 UltraNVIDIA · Open weightArtificial Analysis τ²-bench
83.3%
67
GPT-5.1-CodexOpenAI · ClosedArtificial Analysis τ²-bench
83%
68
GPT-5.1-Codex-MaxOpenAI · ClosedArtificial Analysis τ²-bench
83%
69
GPT-5.1OpenAI · ClosedArtificial Analysis τ²-bench
81.9%
70
o3OpenAI · ClosedArtificial Analysis τ²-bench
80.7%
71
LLaDA2.2-flashInclusionAI · Open weightτ²-bench published setup
80.3%
72
Ternary Bonsai 2 27BPrism ML · Open weightτ²-bench Telecom
80.2%
73
Claude Sonnet 4.6Anthropic · ClosedArtificial Analysis τ²-bench
79.5%
74
DeepSeek V3.2DeepSeek · Open weightArtificial Analysis τ²-bench
78.9%
75
Agents-A1-4BInternScience · Open weightτ²-bench published setup
78.2%
76
GLM-4.6Z.AI · Open weightArtificial Analysis τ²-bench
76.9%
77
Grok Code Fast 1xAI · ClosedArtificial Analysis τ²-bench
75.7%
78
Grok 4xAI · ClosedArtificial Analysis τ²-bench
74.9%
79
K-ExaoneLG AI Research · ClosedArtificial Analysis τ²-bench
74.3%
80
Qwen3 MaxAlibaba · ClosedArtificial Analysis τ²-bench
74.3%
81
Claude Opus 4.7Anthropic · ClosedArtificial Analysis τ²-bench
74%
82
Claude 4.1 Opus ThinkingAnthropic · ClosedArtificial Analysis τ²-bench
71.4%
83
Mercury 2Inception · ClosedArtificial Analysis τ²-bench
70.8%
84
GPT-5 miniOpenAI · ClosedArtificial Analysis τ²-bench
68.4%
85
Nemotron 3 Super 100BNVIDIA · Open weightArtificial Analysis τ²-bench
67.8%
86
Nemotron 3 Super 120B A12BNVIDIA · Open weightArtificial Analysis τ²-bench
67.8%
87
Grok 4 Fast (Reasoning)xAI · ClosedArtificial Analysis τ²-bench
65.8%
88
GPT-OSS 120BOpenAI · Open weightArtificial Analysis τ²-bench
65.8%
89
Grok 4.1 FastxAI · ClosedArtificial Analysis τ²-bench
63.7%
90
o1OpenAI · ClosedArtificial Analysis τ²-bench
62.6%
91
Kimi K2Moonshot AI · ClosedArtificial Analysis τ²-bench
61.1%
92
GPT-OSS 20BOpenAI · Open weightArtificial Analysis τ²-bench
60.2%
93
Gemma 4 31BGoogle · Open weightArtificial Analysis τ²-bench
59.9%
94
LLaDA2.2-miniInclusionAI · Open weightτ²-bench published setup
57.5%
95
Gemini 2.5 ProGoogle · ClosedArtificial Analysis τ²-bench
54.1%
96
GPT-4.1 miniOpenAI · ClosedArtificial Analysis τ²-bench
52.9%
97
Claude 4 SonnetAnthropic · ClosedArtificial Analysis τ²-bench
52.3%
98
DeepSeek V3 0324DeepSeek · Open weightArtificial Analysis τ²-bench
47.1%
99
GPT-4.1OpenAI · ClosedArtificial Analysis τ²-bench
47.1%
100
Sarvam 105BSarvam · Open weightArtificial Analysis τ²-bench
46.8%
101
GLM-4.5-AirZ.AI · ClosedArtificial Analysis τ²-bench
46.5%
102
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weightτ²-bench Telecom
45.3%
103
Gemma 4 26B A4BGoogle · Open weightArtificial Analysis τ²-bench
43.6%
104
Gemini 3 FlashGoogle · ClosedArtificial Analysis τ²-bench
43.3%
105
Mistral Small 4Mistral · Open weightArtificial Analysis τ²-bench
41.2%
106
Mistral Small 4 (Reasoning)Mistral · Open weightArtificial Analysis τ²-bench
41.2%
107
Nemotron 3 Nano 30BNVIDIA · Open weightArtificial Analysis τ²-bench
40.9%
108
DeepSeek V3.1 (Reasoning)DeepSeek · Open weightArtificial Analysis τ²-bench
37.4%
109
North Mini CodeCohere · Open weightArtificial Analysis τ²-bench
37.4%
110
DeepSeek-R1DeepSeek · Open weightArtificial Analysis τ²-bench
36.5%
111
GPT-5 nanoOpenAI · ClosedArtificial Analysis τ²-bench
36.5%
112
Gemma 4 12BGoogle · Open weightArtificial Analysis τ²-bench
36.3%
113
DeepSeek V3.1DeepSeek · Open weightArtificial Analysis τ²-bench
34.8%
114
Sarvam 30BSarvam · Open weightArtificial Analysis τ²-bench
34.5%
115
MiniMax M1 80kMiniMax · ClosedArtificial Analysis τ²-bench
34.2%
116
Solar Pro 2Upstage · ClosedArtificial Analysis τ²-bench
31.9%
117
Mistral Large 2Mistral · ClosedArtificial Analysis τ²-bench
30.7%
118
o3-miniOpenAI · ClosedArtificial Analysis τ²-bench
28.7%
119
Ministral 3 14B (Reasoning)Mistral · Open weightArtificial Analysis τ²-bench
27.2%
120
Ministral 3 14BMistral · Open weightArtificial Analysis τ²-bench
27.2%
121
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weightArtificial Analysis τ²-bench
26.6%
122
Ministral 3 8B (Reasoning)Mistral · Open weightArtificial Analysis τ²-bench
26.6%
123
Ministral 3 8BMistral · Open weightArtificial Analysis τ²-bench
26.6%
124
GPT-4oOpenAI · ClosedArtificial Analysis τ²-bench
25.1%
125
Ministral 3 3B (Reasoning)Mistral · Open weightArtificial Analysis τ²-bench
24.9%
126
Ministral 3 3BMistral · Open weightArtificial Analysis τ²-bench
24.9%
127
Mistral Large 3Mistral · ClosedArtificial Analysis τ²-bench
24.6%
128
Mistral Medium 3Mistral · ClosedArtificial Analysis τ²-bench
24.3%
129
DeepSeek V3DeepSeek · Open weightArtificial Analysis τ²-bench
22.8%
130
Qwen3-Omni-30B-A3B-ThinkingAlibaba · Open weightArtificial Analysis τ²-bench
21.3%
131
Claude 3 HaikuAnthropic · ClosedArtificial Analysis τ²-bench
21.1%
132
Gemma 4 E4BGoogle · Open weightArtificial Analysis τ²-bench
20.8%
133
Gemma 4 E2BGoogle · Open weightArtificial Analysis τ²-bench
20.8%
134
Exaone 4.0 1.2BLG AI Research · Open weightArtificial Analysis τ²-bench
20.5%
135
Granite-4.0-H-1BIBM · Open weightArtificial Analysis τ²-bench
19.6%
136
LFM2.5-1.2B-ThinkingLiquidAI · ClosedArtificial Analysis τ²-bench
19.6%
137
Llama 3.1 405BMeta · Open weightArtificial Analysis τ²-bench
19%
138
Llama 4 MaverickMeta · Open weightArtificial Analysis τ²-bench
17.8%
139
GPT-4.1 nanoOpenAI · ClosedArtificial Analysis τ²-bench
17.3%
140
Qwen3-Omni-30B-A3B-InstructAlibaba · Open weightArtificial Analysis τ²-bench
16.4%
141
Llama 4 ScoutMeta · Open weightArtificial Analysis τ²-bench
15.5%
142
Gemini 2.5 FlashGoogle · ClosedArtificial Analysis τ²-bench
14.9%
143
Granite-4.0-H-350MIBM · Open weightArtificial Analysis τ²-bench
14.6%
144
Nova ProAmazon · ClosedArtificial Analysis τ²-bench
14%
145
Granite-4.0-350MIBM · Open weightArtificial Analysis τ²-bench
13.2%
146
Nemotron Ultra 253BNVIDIA · Open weightArtificial Analysis τ²-bench
11.4%
147
LFM2-24B-A2BLiquidAI · ClosedArtificial Analysis τ²-bench
11.1%
148
LFM2.5-1.2B-InstructLiquidAI · ClosedArtificial Analysis τ²-bench
10.8%
149
Gemma 3 27BGoogle · Open weightArtificial Analysis τ²-bench
10.5%
150
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weightArtificial Analysis τ²-bench
8.5%
151
Exaone 4.0 32BLG AI Research · Open weightArtificial Analysis τ²-bench
4.1%
152
Phi-4Microsoft · Open weightArtificial Analysis τ²-bench
0%

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Use a row only with its attached source and setup label. Match the domain, task release, agent model, user-simulator model, scaffold, prompts, trial count, and pass^k metric before comparing scores. The sorted table is a source ledger, not a controlled cross-provider ranking.

Operator receipt: 152 sourced rows are currently displayable on this page; the highest published score among these 152 models is GLM-5.2 at 99.1%, which does not establish a market leader.

Honest limit: The page mixes a large third-party telecom snapshot with smaller provider-published slices. Those sources do not use one guaranteed-common harness or reporting policy, and BenchLM did not rerun them. A higher number can reflect a different user model, prompt, domain, task release, or repeat policy.

Among the reported τ²-bench results rows, GLM-5.2 is first at 99.1%. The third row is 0.3 points behind. The broader top-10 range is 1.4 points, so many of the published results sit in a relatively narrow band.

152 models have been evaluated on τ²-bench results. The benchmark falls in the Agentic category. τ²-bench results is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About τ²-bench results

Year

2025

Tasks

Airline, retail, and telecom customer-service task sets

Format

Published domain success or pass^k results

Difficulty

Dual-control customer-service workflows

τ²-bench extends the original benchmark with a dual-control telecom domain where the agent and simulated user can both act through tools. The maintained framework also includes airline and retail. BenchLM keeps each exact source attached and labels the published setup instead of treating every row as one controlled run.

Freshness and provenance

Version

τ²-Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

Are all τ²-bench scores directly comparable?

No. Match the domain, task release, agent model, user model, scaffold, prompts, number of trials, and pass^k definition. BenchLM labels each sourced row so a telecom result or third-party implementation is not silently treated as the same setup as an airline, retail, or aggregate result.

Does a high τ²-bench score prove production support reliability?

No. It measures success in simulated customer-service environments under a reported setup. Production identity checks, permission boundaries, changing policies, latency, cost, monitoring, and human escalation still need separate testing.

Last updated: October 2, 2026 · BenchLM version τ²-Bench 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.