Skip to main content
BenchLM

Vals Tax Agent Bench (Tax Agent Bench)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Agent evaluation on research-grade US tax questions using the Tax Agent Bench harness.

Overall score on Tax Agent Bench — September 22, 2026

We mirror the published overall score view for Tax Agent Bench. Claude Fable 5.1 leads the public snapshot at 77.64%, followed by Claude Opus 5 (75.06%) and GLM-5.3 (73.09%). We do not use these results to rank models overall.

26 modelsAgenticCurrentDisplay onlyUpdated September 22, 2026

Overall score table (26 models)

Score
1
Claude Fable 5.1Anthropic · Closed
77.64%
2
Claude Opus 5Anthropic · Closed
75.06%
3
GLM-5.3Z.AI · Open weightmax reasoning
73.09%
4
Muse Spark 1.3Meta · Closedxhigh reasoning
71.93%
5
Grok 4.6xAI · Closedhigh reasoning
70.79%
6
Claude Opus 5.5Anthropic · Closed
70.50%
7
Kimi K3Moonshot AI · Closedmax reasoning
68.67%
8
GPT-5.6 SolOpenAI · Closedmax reasoning
67.95%
9
Gemini 3.8 FlashGoogle · Closedhigh reasoning
66.77%
10
Qwen3.8 MaxAlibaba · Open weight
65.96%
11
Grok 4.7xAI · Closedxhigh reasoning
65.60%
12
GPT-5.6 TerraOpenAI · Closedmax reasoning
65.20%
13
Hy4 previewTencent · Open weight
63.71%
14
GPT-6 AstraOpenAI · Closedmax reasoning
63.34%
15
DeepSeek V4.1 FlashDeepSeek · Open weighthigh reasoning
62.46%
16
Claude Sonnet 5Anthropic · Closed
62.27%
17
GPT-5.6 LunaOpenAI · Closedmax reasoning
60.81%
18
GPT-5.5OpenAI · Closedxhigh reasoning
60.46%
19
GPT-6 LunaOpenAI · Closedmax reasoning
58.86%
20
DeepSeek V4 Pro 0813DeepSeek · Open weightmax reasoning
58.66%
21
Gemini 3.7 FlashGoogle · Closedhigh reasoning
57.66%
22
GPT-6 SolOpenAI · Closedmax reasoning
53.05%
23
MiniMax M3MiniMax · Open weight
49.69%
25
InklingThinking Machines Lab · Open weight0.99 reasoning
43.52%
26
Mercury 2.5Inception · Closedhigh reasoning
12.77%

How Tax Agent Bench is shown here

BenchLM mirrors the public Vals AI Tax Agent Bench leaderboard captured from https://www.vals.ai/benchmarks/tax_agent_bench and updated by Vals on September 22, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Tax Agent Bench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

26 Vals rows14 task viewsprivate datasetTasks: Overall, All-Pass, Fact-Pattern Analysis, Rule & Source Lookup, Current & Temporal AnalysisDisplay only

The published Tax Agent Bench snapshot places Claude Fable 5.1 first at 77.64%. The third row is 4.55 points behind. The broader top-10 range is 11.68 points, so the table still separates the published systems.

26 models have been evaluated on Tax Agent Bench. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Tax Agent Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Tax Agent Bench

Year

2026

Tasks

US tax research, calculations, forms, filings, and precedent analysis

Format

Overall score; all-pass score shown separately

Difficulty

Research-grade US tax questions

The public table preserves overall and all-pass scores across fact-pattern analysis, source lookup, temporal analysis, calculations, filings, and precedent. These private-dataset agent results remain display-only and do not enter weighted model rankings.

Freshness and provenance

Version

Tax Agent Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Tax Agent Bench measure?

Agent evaluation on research-grade US tax questions using the Tax Agent Bench harness.

Which model leads the published Tax Agent Bench snapshot?

Claude Fable 5.1 currently leads the published Tax Agent Bench snapshot with 77.64% overall score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Tax Agent Bench?

The September 22, 2026 snapshot contains 26 AI models.

Last updated: September 22, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.