Skip to main content
BenchLM

Vals TaxEval v2 (TaxEval v2)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A Vals-created set of questions and responses to tax questions

TaxEval v2 score on TaxEval v2 — September 1, 2026

We mirror the published taxeval v2 score view for TaxEval v2. Muse Spark 1.2 leads the public snapshot at 80.38%, followed by Muse Spark 1.1 (79.72%) and Muse Spark (77.68%). We do not use these results to rank models overall.

145 modelsAgenticCurrentDisplay onlyUpdated September 1, 2026

TaxEval v2 score table (145 models)

Score
1
Muse Spark 1.2Meta · Closedxhigh reasoning
80.38%
2
Muse Spark 1.1Meta · Closedxhigh reasoning
79.72%
3
Muse SparkMeta · Closed
77.68%
4
Claude Sonnet 4.6Anthropic · Closed
77.11%
5
Claude Fable 5Anthropic · Closed
76.94%
6
GPT-5.6 TerraOpenAI · Closedxhigh reasoning
76.17%
7
GPT-5.6 LunaOpenAI · Closedmax reasoning
76.17%
8
Claude Opus 4.6 (Adaptive)Anthropic · Closed
75.96%
9
Claude Fable 5.1Anthropic · Closed
75.96%
10
75.88%
11
GPT-5.2OpenAI · Closedxhigh reasoning
75.76%
12
Kimi K3Moonshot AI · Closed
75.72%
13
75.70%
14
Claude Opus 4.8Anthropic · Closed
75.63%
15
Claude Sonnet 5Anthropic · Closed
75.63%
16
GLM-5.3-FlashZ.AI · Open weightmax reasoning
75.59%
17
Qwen3.8 MaxAlibaba · Open weight
75.55%
18
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
75.51%
19
Qwen3.7 MaxAlibaba · Closed
75.31%
20
InklingThinking Machines Lab · Open weight0.99 reasoning
75.31%
21
Claude Opus 4.7Anthropic · Closed
75.27%
22
GPT-5 miniOpenAI · Closedhigh reasoning
75.22%
23
Claude Opus 5Anthropic · Closed
75.14%
24
GPT-4.1OpenAI · Closedhigh reasoning
75.06%
25
GPT-5.5OpenAI · Closedxhigh reasoning
74.98%
26
Gemini 3.6 FlashGoogle · Closedhigh reasoning
74.86%
27
GPT-5.1OpenAI · Closedhigh reasoning
74.86%
28
Claude Opus 4.5 ThinkingAnthropic · Closed
74.86%
29
O4 MiniOpenAIhigh reasoning
74.78%
30
GPT-5.6 SolOpenAI · Closedmax reasoning
74.78%
31
Gemini 3.7 FlashGoogle · Closedhigh reasoning
74.73%
32
Qwen3.6 PlusAlibaba · Closed
74.73%
33
Kimi K2.6Moonshot AI · Open weight
74.65%
34
o3OpenAI · Closedhigh reasoning
74.57%
35
GPT-4oOpenAI · Closedhigh reasoning
74.53%
36
Gemini 3.8 FlashGoogle · Closedhigh reasoning
74.45%
37
Gemini 3.5 FlashGoogle · Closedhigh reasoning
74.37%
38
Claude Opus 4.5Anthropic · Closed
74.33%
39
o1OpenAI · Closedhigh reasoning
74.28%
40
74.20%
43
73.96%
44
GPT-5.4OpenAI · Closedxhigh reasoning
73.96%
45
Gemini 3 Flash PreviewGooglehigh reasoning
73.88%
46
MiMo-V2.5-ProXiaomi · Closed
73.79%
48
Qwen3 MaxAlibaba · Closed
73.51%
49
GPT-5OpenAIhigh reasoning
73.39%
50
GLM-5.2Z.AI · Open weight
73.34%
51
Claude Sonnet 4.5 ThinkingAnthropic · Closed
73.30%
52
73.14%
54
73.06%
55
DeepSeek V4 Pro 0813DeepSeek · Open weightmax reasoning
73.06%
56
72.98%
58
Gemini 3.1 Pro PreviewGooglehigh reasoning
72.88%
60
MiniMax M3MiniMax · Open weight
72.73%
61
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
72.61%
62
Gemini 3 Pro PreviewGooglehigh reasoning
72.57%
63
72.40%
65
GLM-4.5Z.AI · Closed
72.40%
66
GLM-5.3Z.AI · Open weightmax reasoning
72.36%
67
DeepSeek-R1DeepSeek · Open weight
72.28%
68
Qwen3.5 FlashAlibaba · Closed
72.16%
69
DeepSeek V4 Pro 0813DeepSeek · Open weightmax reasoning
72.08%
71
GPT-4.1 miniOpenAI · Closedhigh reasoning
71.91%
72
Claude Opus 4Anthropic
71.91%
73
MiMo-V2.5Xiaomi · Closed
71.83%
74
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
71.79%
75
Kimi K2 ThinkingMoonshot AI
71.71%
76
Grok 4.5xAI · Closedhigh reasoning
71.67%
77
GPT-OSS 120BOpenAI · Open weight
71.59%
79
71.46%
80
Qwen3.6-27BAlibaba · Open weight
71.26%
81
GPT-5.4 miniOpenAI · Closedxhigh reasoning
71.22%
82
GLM-5.1Z.AI · Open weight
71.19%
84
GPT-4oOpenAI · Closedhigh reasoning
71.14%
85
71.14%
86
DeepSeek V3 0324Fireworks AI
71.10%
87
Grok 4.6xAI · Closedhigh reasoning
71.10%
88
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
70.85%
89
Grok 4.3xAI · Closed
70.81%
90
DeepSeek V4 Flash 0731DeepSeek · Open weighthigh reasoning
70.69%
92
Qwen3 235b A22bFireworks AI
70.65%
94
70.32%
95
70.20%
96
Claude 3.5 SonnetAnthropic · Closed
70.16%
97
70.03%
99
69.62%
100
o3-miniOpenAI · Closedhigh reasoning
69.42%
101
GLM-4.7Z.AI · Open weight
68.77%
103
MiniMax M2.5MiniMax · Closed
68.15%
104
DeepSeek V3p2 ThinkingFireworks AIhigh reasoning
68.15%
105
Mistral Medium 3.5Mistral AIhigh reasoning
67.99%
106
DeepSeek V3DeepSeek · Open weight
67.91%
107
67.74%
108
Claude Haiku 4.5 ThinkingAnthropic · Closed
67.54%
109
GPT-5.4 nanoOpenAI · Closedhigh reasoning
67.42%
110
GPT-5 nanoOpenAI · Closedhigh reasoning
67.38%
111
67.05%
112
66.56%
113
MiniMax M2.7MiniMax · Open weight
66.56%
114
66.35%
115
GLM-4.6Z.AI · Open weight
66.23%
117
65.25%
118
65.09%
120
63.78%
122
GPT-OSS 20BOpenAI · Open weight
63.70%
123
61.94%
125
61.37%
126
60.88%
128
GPT-4.1 nanoOpenAI · Closedhigh reasoning
60.75%
129
GPT-4o miniOpenAI · Closedhigh reasoning
60.55%
130
60.30%
132
59.48%
134
Laguna XS.2Poolside · Open weight
58.95%
135
58.30%
136
58.18%
137
57.36%
140
49.14%
141
48.20%
142
44.60%
143
41.86%
145
Laguna M.1Poolside · Closed
1.64%

How TaxEval v2 is shown here

BenchLM mirrors the public Vals AI TaxEval v2 leaderboard captured from https://www.vals.ai/benchmarks/tax_eval_v2 and updated by Vals on September 1, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

TaxEval v2 is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

145 Vals rows3 task viewsprivate datasetTasks: Overall, Correctness, Stepwise ReasoningDisplay only

The published TaxEval v2 snapshot places Muse Spark 1.2 first at 80.38%. The third row is 2.70 points behind. The broader top-10 range is 4.50 points, so many of the published results sit in a relatively narrow band.

145 models have been evaluated on TaxEval v2. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. TaxEval v2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About TaxEval v2

Year

2026

Tasks

Tax question answering and response evaluation

Format

Accuracy score

Difficulty

Professional tax reasoning

BenchLM mirrors the public Vals AI TaxEval v2 leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

Freshness and provenance

Version

TaxEval v2 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does TaxEval v2 measure?

A Vals-created set of questions and responses to tax questions

Which model leads the published TaxEval v2 snapshot?

Muse Spark 1.2 currently leads the published TaxEval v2 snapshot with 80.38% taxeval v2 score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on TaxEval v2?

The September 1, 2026 snapshot contains 145 AI models.

Last updated: September 1, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.