Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

τ²-Bench Tool-Agent-User Evaluation (τ²-bench results)

This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Use a row only with its attached source and setup label. Match the domain, task release, agent model, user-simulator model, scaffold, prompts, trial count, and pass^k metric before comparing scores. The sorted table is a source ledger, not a controlled cross-provider ranking.

Operator receipt: 151 sourced rows are currently displayable on this page; the highest published score among these 151 models is GLM-5.2 at 99.1%, which does not establish a market leader.

Honest limit: The page mixes a large third-party telecom snapshot with smaller provider-published slices. Those sources do not use one guaranteed-common harness or reporting policy, and BenchLM did not rerun them. A higher number can reflect a different user model, prompt, domain, task release, or repeat policy.

Benchmark score on τ²-bench sourced results — September 10, 2026

We mirror the published score view for τ²-bench sourced results. The public snapshot contains 151 models. GLM-5.2 has the highest published score at 99.1%, but the available coverage and evaluation setups do not establish a market leader. We do not use these results to rank models overall.

151 modelsAgenticCurrentDisplay onlyUpdated September 10, 2026

Benchmark score table (151 models)

Score
1
GLM-5.2Z.AI · Open weightArtificial Analysis τ²-bench
99.1%
2
GPT-5.4OpenAI · Closedτ²-bench Telecom
98.9%
3
GLM-4.7-FlashZ.AI · Open weightArtificial Analysis τ²-bench
98.8%
4
Claude Fable 5Anthropic · ClosedArtificial Analysis τ²-bench
98.5%
5
GLM-5-TurboZ.AI · ClosedArtificial Analysis τ²-bench
98.5%
6
GLM-5V-TurboZ.AI · ClosedArtificial Analysis τ²-bench
98.5%
7
Step 3.7 FlashStepFun · Open weightArtificial Analysis τ²-bench
98.5%
8
GLM-5Z.AI · Open weightArtificial Analysis τ²-bench
98.2%
9
GPT-5.5OpenAI · Closedτ²-bench Telecom
98%
10
GLM-5.1Z.AI · Open weightArtificial Analysis τ²-bench
97.7%
11
Qwen3.6 PlusAlibaba · ClosedArtificial Analysis τ²-bench
97.7%
12
Grok 4.3xAI · ClosedArtificial Analysis τ²-bench
97.7%
13
DeepSeek V4 Pro 0813DeepSeek · ClosedArtificial Analysis τ²-bench
96.2%
14
Kimi K2.6Moonshot AI · Open weightArtificial Analysis τ²-bench
95.9%
15
Kimi K2.5Moonshot AI · Open weightArtificial Analysis τ²-bench
95.9%
16
GLM-4.7Z.AI · Open weightArtificial Analysis τ²-bench
95.9%
17
Qwen 3.6 Max (preview)Alibaba · ClosedArtificial Analysis τ²-bench
95.9%
18
Kimi K2.5 (Reasoning)Moonshot AI · ClosedArtificial Analysis τ²-bench
95.9%
19
Gemini 3.1 ProGoogle · Closedτ²-bench published setup
95.6%
20
Gemini 3.5 FlashGoogle · ClosedArtificial Analysis τ²-bench
95.3%
21
Qwen3.6-35B-A3BAlibaba · Open weightArtificial Analysis τ²-bench
95.3%
22
MiniMax M2.5MiniMax · ClosedArtificial Analysis τ²-bench
95.3%
23
MiMo-V2-ProXiaomi · ClosedArtificial Analysis τ²-bench
95%
24
Qwen3.7 MaxAlibaba · ClosedArtificial Analysis τ²-bench
94.7%
25
Claude Opus 4.8Anthropic · ClosedArtificial Analysis τ²-bench
94.4%
26
Mistral Medium 3.5 128BMistral · Open weightArtificial Analysis τ²-bench
94.2%
27
Qwen3.6-27BAlibaba · Open weightArtificial Analysis τ²-bench
94.2%
28
MiMo-V2.5-ProXiaomi · ClosedArtificial Analysis τ²-bench
94.2%
29
DeepSeek V4 Pro (High)DeepSeek · Open weightArtificial Analysis τ²-bench
94.2%
30
Qwen3.5-27BAlibaba · Open weightArtificial Analysis τ²-bench
93.9%
31
Qwen3.5-122B-A10BAlibaba · Open weightArtificial Analysis τ²-bench
93.6%
32
GPT-5.4 miniOpenAI · Closedτ²-bench Telecom
93.4%
33
Grok 4.1 Fast (Reasoning)xAI · ClosedArtificial Analysis τ²-bench
93.3%
34
Qwen3.7 PlusAlibaba · ClosedArtificial Analysis τ²-bench
93%
35
GPT-5.4 nanoOpenAI · Closedτ²-bench Telecom
92.5%
36
GPT-5.2-CodexOpenAI · ClosedArtificial Analysis τ²-bench
92.1%
37
Claude Opus 4.6 (Adaptive)Anthropic · ClosedArtificial Analysis τ²-bench
92.1%
38
Muse SparkMeta · Closedτ²-bench published setup
91.5%
39
MiMo-V2-OmniXiaomi · ClosedArtificial Analysis τ²-bench
91.2%
40
Kimi K2.7 CodeMoonshot AI · Open weightArtificial Analysis τ²-bench
90.1%
41
Trinity-Large-ThinkingArcee AI · Open weightArtificial Analysis τ²-bench
90.1%
42
Trinity-Large-PreviewArcee AI · Open weightArtificial Analysis τ²-bench
90.1%
43
Claude Opus 4.5 ThinkingAnthropic · ClosedArtificial Analysis τ²-bench
89.5%
44
Qwen3.5-35B-A3BAlibaba · Open weightArtificial Analysis τ²-bench
89.2%
45
MiniMax M3MiniMax · Open weightArtificial Analysis τ²-bench
88.9%
46
Claude Opus 4.7 (Adaptive)Anthropic · ClosedArtificial Analysis τ²-bench
88.6%
47
LFM2.5-8B-A1BLiquidAI · Open weightτ²-bench Telecom
88.1%
48
Step 3.5 FlashStepFun · Open weightArtificial Analysis τ²-bench
87.4%
49
Gemini 3 ProGoogle · ClosedArtificial Analysis τ²-bench
87.1%
50
GPT-5 (medium)OpenAI · ClosedArtificial Analysis τ²-bench
86.5%
51
GPT-5.6 TerraOpenAI · ClosedArtificial Analysis τ²-bench
86.3%
52
Claude Opus 4.5Anthropic · ClosedArtificial Analysis τ²-bench
86.3%
53
Solar Pro 3Upstage · ClosedArtificial Analysis τ²-bench
86.3%
54
GPT-5.3 CodexOpenAI · ClosedArtificial Analysis τ²-bench
86%
55
Ling 2.6 FlashInclusionAI · Open weightArtificial Analysis τ²-bench
86%
56
GPT-5.3-Codex-SparkOpenAI · ClosedArtificial Analysis τ²-bench
86%
57
GPT-5.6 SolOpenAI · ClosedArtificial Analysis τ²-bench
85.1%
58
Command A+Cohere · Open weightτ²-bench Telecom
85%
59
Claude Opus 4.6Anthropic · ClosedArtificial Analysis τ²-bench
84.8%
60
MiniMax M2.7MiniMax · Open weightArtificial Analysis τ²-bench
84.8%
61
GPT-5.2OpenAI · ClosedArtificial Analysis τ²-bench
84.8%
62
GPT-5 (high)OpenAI · ClosedArtificial Analysis τ²-bench
84.8%
63
Qwen3.5 397BAlibaba · Open weightArtificial Analysis τ²-bench
83.9%
64
MiMo-V2-FlashXiaomi · Open weightArtificial Analysis τ²-bench
83.9%
65
Qwen3.5 397B (Reasoning)Alibaba · Open weightArtificial Analysis τ²-bench
83.9%
66
Nemotron 3 UltraNVIDIA · Open weightArtificial Analysis τ²-bench
83.3%
67
GPT-5.1-CodexOpenAI · ClosedArtificial Analysis τ²-bench
83%
68
GPT-5.1-Codex-MaxOpenAI · ClosedArtificial Analysis τ²-bench
83%
69
GPT-5.1OpenAI · ClosedArtificial Analysis τ²-bench
81.9%
70
o3OpenAI · ClosedArtificial Analysis τ²-bench
80.7%
71
LLaDA2.2-flashInclusionAI · Open weightτ²-bench published setup
80.3%
72
Claude Sonnet 4.6Anthropic · ClosedArtificial Analysis τ²-bench
79.5%
73
DeepSeek V3.2DeepSeek · Open weightArtificial Analysis τ²-bench
78.9%
74
Agents-A1-4BInternScience · Open weightτ²-bench published setup
78.2%
75
GLM-4.6Z.AI · Open weightArtificial Analysis τ²-bench
76.9%
76
Grok Code Fast 1xAI · ClosedArtificial Analysis τ²-bench
75.7%
77
Grok 4xAI · ClosedArtificial Analysis τ²-bench
74.9%
78
K-ExaoneLG AI Research · ClosedArtificial Analysis τ²-bench
74.3%
79
Qwen3 MaxAlibaba · ClosedArtificial Analysis τ²-bench
74.3%
80
Claude Opus 4.7Anthropic · ClosedArtificial Analysis τ²-bench
74%
81
Claude 4.1 Opus ThinkingAnthropic · ClosedArtificial Analysis τ²-bench
71.4%
82
Mercury 2Inception · ClosedArtificial Analysis τ²-bench
70.8%
83
GPT-5 miniOpenAI · ClosedArtificial Analysis τ²-bench
68.4%
84
Nemotron 3 Super 100BNVIDIA · Open weightArtificial Analysis τ²-bench
67.8%
85
Nemotron 3 Super 120B A12BNVIDIA · Open weightArtificial Analysis τ²-bench
67.8%
86
GPT-OSS 120BOpenAI · Open weightArtificial Analysis τ²-bench
65.8%
87
Grok 4 Fast (Reasoning)xAI · ClosedArtificial Analysis τ²-bench
65.8%
88
Grok 4.1 FastxAI · ClosedArtificial Analysis τ²-bench
63.7%
89
o1OpenAI · ClosedArtificial Analysis τ²-bench
62.6%
90
Kimi K2Moonshot AI · ClosedArtificial Analysis τ²-bench
61.1%
91
GPT-OSS 20BOpenAI · Open weightArtificial Analysis τ²-bench
60.2%
92
Gemma 4 31BGoogle · Open weightArtificial Analysis τ²-bench
59.9%
93
LLaDA2.2-miniInclusionAI · Open weightτ²-bench published setup
57.5%
94
Gemini 2.5 ProGoogle · ClosedArtificial Analysis τ²-bench
54.1%
95
GPT-4.1 miniOpenAI · ClosedArtificial Analysis τ²-bench
52.9%
96
Claude 4 SonnetAnthropic · ClosedArtificial Analysis τ²-bench
52.3%
97
GPT-4.1OpenAI · ClosedArtificial Analysis τ²-bench
47.1%
98
Sarvam 105BSarvam · Open weightArtificial Analysis τ²-bench
46.8%
99
GLM-4.5-AirZ.AI · ClosedArtificial Analysis τ²-bench
46.5%
100
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weightτ²-bench Telecom
45.3%
101
Gemma 4 26B A4BGoogle · Open weightArtificial Analysis τ²-bench
43.6%
102
Gemini 3 FlashGoogle · ClosedArtificial Analysis τ²-bench
43.3%
103
Mistral Small 4Mistral · Open weightArtificial Analysis τ²-bench
41.2%
104
Mistral Small 4 (Reasoning)Mistral · Open weightArtificial Analysis τ²-bench
41.2%
105
Nemotron 3 Nano 30BNVIDIA · Open weightArtificial Analysis τ²-bench
40.9%
106
DeepSeek V3.1 (Reasoning)DeepSeek · Open weightArtificial Analysis τ²-bench
37.4%
107
North Mini CodeCohere · Open weightArtificial Analysis τ²-bench
37.4%
108
DeepSeek-R1DeepSeek · Open weightArtificial Analysis τ²-bench
36.5%
109
GPT-5 nanoOpenAI · ClosedArtificial Analysis τ²-bench
36.5%
110
Gemma 4 12BGoogle · Open weightArtificial Analysis τ²-bench
36.3%
111
DeepSeek V3.1DeepSeek · Open weightArtificial Analysis τ²-bench
34.8%
112
Sarvam 30BSarvam · Open weightArtificial Analysis τ²-bench
34.5%
113
MiniMax M1 80kMiniMax · ClosedArtificial Analysis τ²-bench
34.2%
114
Solar Pro 2Upstage · ClosedArtificial Analysis τ²-bench
31.9%
115
Mistral Large 2Mistral · ClosedArtificial Analysis τ²-bench
30.7%
116
o3-miniOpenAI · ClosedArtificial Analysis τ²-bench
28.7%
117
Ministral 3 14B (Reasoning)Mistral · Open weightArtificial Analysis τ²-bench
27.2%
118
Ministral 3 14BMistral · Open weightArtificial Analysis τ²-bench
27.2%
119
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weightArtificial Analysis τ²-bench
26.6%
120
Ministral 3 8B (Reasoning)Mistral · Open weightArtificial Analysis τ²-bench
26.6%
121
Ministral 3 8BMistral · Open weightArtificial Analysis τ²-bench
26.6%
122
GPT-4oOpenAI · ClosedArtificial Analysis τ²-bench
25.1%
123
Ministral 3 3B (Reasoning)Mistral · Open weightArtificial Analysis τ²-bench
24.9%
124
Ministral 3 3BMistral · Open weightArtificial Analysis τ²-bench
24.9%
125
Mistral Large 3Mistral · ClosedArtificial Analysis τ²-bench
24.6%
126
Mistral Medium 3Mistral · ClosedArtificial Analysis τ²-bench
24.3%
127
DeepSeek V3DeepSeek · Open weightArtificial Analysis τ²-bench
22.8%
128
Granite-4.0-1BIBM · Open weightArtificial Analysis τ²-bench
22.8%
129
Qwen3-Omni-30B-A3B-ThinkingAlibaba · Open weightArtificial Analysis τ²-bench
21.3%
130
Claude 3 HaikuAnthropic · ClosedArtificial Analysis τ²-bench
21.1%
131
Gemma 4 E4BGoogle · Open weightArtificial Analysis τ²-bench
20.8%
132
Gemma 4 E2BGoogle · Open weightArtificial Analysis τ²-bench
20.8%
133
Exaone 4.0 1.2BLG AI Research · Open weightArtificial Analysis τ²-bench
20.5%
134
Granite-4.0-H-1BIBM · Open weightArtificial Analysis τ²-bench
19.6%
135
LFM2.5-1.2B-ThinkingLiquidAI · ClosedArtificial Analysis τ²-bench
19.6%
136
Llama 3.1 405BMeta · Open weightArtificial Analysis τ²-bench
19%
137
Llama 4 MaverickMeta · Open weightArtificial Analysis τ²-bench
17.8%
138
GPT-4.1 nanoOpenAI · ClosedArtificial Analysis τ²-bench
17.3%
139
Qwen3-Omni-30B-A3B-InstructAlibaba · Open weightArtificial Analysis τ²-bench
16.4%
140
Llama 4 ScoutMeta · Open weightArtificial Analysis τ²-bench
15.5%
141
Gemini 2.5 FlashGoogle · ClosedArtificial Analysis τ²-bench
14.9%
142
Granite-4.0-H-350MIBM · Open weightArtificial Analysis τ²-bench
14.6%
143
Nova ProAmazon · ClosedArtificial Analysis τ²-bench
14%
144
Granite-4.0-350MIBM · Open weightArtificial Analysis τ²-bench
13.2%
145
Nemotron Ultra 253BNVIDIA · Open weightArtificial Analysis τ²-bench
11.4%
146
LFM2-24B-A2BLiquidAI · ClosedArtificial Analysis τ²-bench
11.1%
147
LFM2.5-1.2B-InstructLiquidAI · ClosedArtificial Analysis τ²-bench
10.8%
148
Gemma 3 27BGoogle · Open weightArtificial Analysis τ²-bench
10.5%
149
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weightArtificial Analysis τ²-bench
8.5%
150
Exaone 4.0 32BLG AI Research · Open weightArtificial Analysis τ²-bench
4.1%
151
Phi-4Microsoft · Open weightArtificial Analysis τ²-bench
0%

The published τ²-bench results snapshot places GLM-5.2 first at 99.1%. The third row is 0.3 points behind. The broader top-10 range is 1.4 points, so many of the published results sit in a relatively narrow band.

151 models have been evaluated on τ²-bench results. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. τ²-bench results is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About τ²-bench results

Year

2025

Tasks

Airline, retail, and telecom customer-service task sets

Format

Published domain success or pass^k results

Difficulty

Dual-control customer-service workflows

τ²-bench extends the original benchmark with a dual-control telecom domain where the agent and simulated user can both act through tools. The maintained framework also includes airline and retail. BenchLM keeps each exact source attached and labels the published setup instead of treating every row as one controlled run.

BenchLM freshness & provenance

Version

τ²-Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

Are all τ²-bench scores directly comparable?

No. Match the domain, task release, agent model, user model, scaffold, prompts, number of trials, and pass^k definition. BenchLM labels each sourced row so a telecom result or third-party implementation is not silently treated as the same setup as an airline, retail, or aggregate result.

Does a high τ²-bench score prove production support reliability?

No. It measures success in simulated customer-service environments under a reported setup. Production identity checks, permission boundaries, changing policies, latency, cost, monitoring, and human escalation still need separate testing.

Last updated: September 10, 2026 · BenchLM version τ²-Bench 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.