Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Artificial Analysis Tau3-Banking (AA Tau3 Banking)

We mirror this table; we do not rank on it.

Data verified 35 confirmed releases in the last 30 daysFollow model changes

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

Benchmark score on AA Tau3 Banking — September 18, 2026

We mirror the published score view for AA Tau3 Banking. Grok 4.6 leads the public snapshot at 50.7%, followed by Muse Spark 1.3 (50.5%) and GLM-5.3 (50.3%). We do not use these results to rank models overall.

15 modelsAgenticCurrentDisplay onlyUpdated September 18, 2026

Benchmark score table (15 models)

Score
1
Grok 4.6xAI · Closed
50.7%
2
Muse Spark 1.3Meta · Closed
50.5%
3
GLM-5.3Z.AI · Open weight
50.3%
4
Qwen3.8-27BAlibaba · Open weight
48.0%
5
Claude Fable 5.1Anthropic · Closed
47.2%
6
GLM-5.3-FlashZ.AI · Open weight
47.2%
7
Kimi K3Moonshot AI · Closed
46.0%
8
Gemini 3.8 FlashGoogle · Closed
44.9%
9
GPT-5.6 SolOpenAI · Closed
44.3%
10
Claude Opus 5Anthropic · Closed
42.1%
11
GPT-6 AstraOpenAI · Closed
41.4%
12
GPT-5.6 TerraOpenAI · Closed
40.2%
13
DeepSeek V4 Pro 0813DeepSeek · Closed
39.6%
14
Claude Fable 5Anthropic · Closed
38.1%
15
Ling 3.0 FlashInclusionAI · Open weight
28.0%

The published AA Tau3 Banking snapshot places Grok 4.6 first at 50.7%. The third row is 0.4 points behind. The broader top-10 range is 8.6 points, so many of the published results sit in a relatively narrow band.

15 models have been evaluated on AA Tau3 Banking. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. AA Tau3 Banking is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA Tau3 Banking

Year

2026

Tasks

Banking tool-use workflows

Format

Task success rate

Difficulty

Agentic banking workflows

Stored separately from the broader Tau3-Bench lane.

Freshness and provenance

Version

AA Tau3 Banking 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does AA Tau3 Banking measure?

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

Which model scores highest on AA Tau3 Banking?

Grok 4.6 by xAI currently leads with a score of 50.7% on AA Tau3 Banking.

How many models are evaluated on AA Tau3 Banking?

15 AI models have been evaluated on AA Tau3 Banking on BenchLM.

Last updated: September 18, 2026 · BenchLM version AA Tau3 Banking 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.