Skip to main content
BenchLM

Vals MortgageTax (MortgageTax)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Vals AI benchmark for mortgage and tax document reasoning, including semantic and numerical extraction task views.

MortgageTax score on MortgageTax — September 1, 2026

We mirror the published mortgagetax score view for MortgageTax. Claude Opus 5 leads the public snapshot at 72.06%, followed by Claude Fable 5.1 (70.79%) and Claude Opus 4.7 (70.27%). We do not use these results to rank models overall.

98 modelsAgenticCurrentDisplay onlyUpdated September 1, 2026

MortgageTax score table (98 models)

Score
1
Claude Opus 5Anthropic · Closed
72.06%
2
Claude Fable 5.1Anthropic · Closed
70.79%
3
Claude Opus 4.7Anthropic · Closed
70.27%
4
Claude Sonnet 5Anthropic · Closed
70.03%
5
Claude Opus 4.8Anthropic · Closed
69.91%
6
Gemini 3.1 Pro PreviewGooglehigh reasoning
69.40%
7
Gemini 3 Pro PreviewGooglehigh reasoning
69.08%
8
Claude Fable 5Anthropic · Closed
68.92%
9
Gemini 2.5 ProGoogle · Closed
68.92%
10
GPT-5.5OpenAI · Closedxhigh reasoning
68.76%
11
Gemini 3 Flash PreviewGooglehigh reasoning
68.72%
12
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
68.68%
13
68.68%
14
Claude Opus 4.5Anthropic · Closed
68.68%
15
Claude Opus 4.6 (Adaptive)Anthropic · Closed
68.52%
16
MiniMax M3MiniMax · Open weight
68.36%
17
GPT-5.4OpenAI · Closedxhigh reasoning
68.32%
18
Qwen3.6-27BAlibaba · Open weight
68.28%
19
Gemini 3.5 FlashGoogle · Closedhigh reasoning
68.12%
20
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
68.04%
21
Qwen3.6 PlusAlibaba · Closed
67.97%
22
Claude Sonnet 4.6Anthropic · Closed
67.73%
23
Claude Opus 4.5 ThinkingAnthropic · Closed
67.69%
24
Gemini 3.6 FlashGoogle · Closedhigh reasoning
67.61%
25
Qwen3.5 FlashAlibaba · Closed
67.37%
26
GPT-5.6 TerraOpenAI · Closedxhigh reasoning
67.33%
27
GPT-5.6 LunaOpenAI · Closedmax reasoning
67.29%
28
GPT-5.6 SolOpenAI · Closedmax reasoning
67.29%
30
GPT-5.2OpenAI · Closedxhigh reasoning
67.13%
31
GPT-5 miniOpenAI · Closedhigh reasoning
66.89%
33
Gemini 3.7 FlashGoogle · Closedhigh reasoning
66.65%
34
66.53%
35
Kimi K3Moonshot AI · Closed
66.34%
36
Qwen3.7 PlusAlibaba · Closed
66.18%
37
Muse Spark 1.1Meta · Closedxhigh reasoning
66.14%
38
GPT-4.1OpenAI · Closedhigh reasoning
65.94%
39
Kimi K2.6Moonshot AI · Open weight
65.82%
40
o3OpenAI · Closedhigh reasoning
65.70%
41
GPT-4.1 miniOpenAI · Closedhigh reasoning
65.50%
42
GPT-5OpenAIhigh reasoning
65.45%
43
Muse Spark 1.2Meta · Closedxhigh reasoning
65.42%
44
Gemini 3.8 FlashGoogle · Closedhigh reasoning
65.34%
45
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
64.94%
46
O4 MiniOpenAIhigh reasoning
64.83%
47
Grok 4.6xAI · Closedhigh reasoning
64.19%
48
Claude 3.5 SonnetAnthropic · Closed
64.07%
49
Qwen3.8 MaxAlibaba · Open weight
63.99%
50
Claude Sonnet 4.5 ThinkingAnthropic · Closed
63.99%
51
InklingThinking Machines Lab · Open weight0.99 reasoning
63.99%
52
GPT-5.4 miniOpenAI · Closedxhigh reasoning
63.51%
54
62.47%
55
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
62.28%
56
Claude Haiku 4.5 ThinkingAnthropic · Closed
62.16%
57
62.12%
59
Grok 4.5xAI · Closedhigh reasoning
61.80%
60
Gemma 4 31b ItGooglehigh reasoning
61.37%
61
GPT-5.1OpenAI · Closedhigh reasoning
61.37%
62
61.21%
63
61.09%
64
GPT-4oOpenAI · Closedhigh reasoning
60.97%
65
60.77%
66
60.37%
67
59.66%
68
MiMo-V2.5Xiaomi · Closed
59.26%
69
GPT-5.4 nanoOpenAI · Closedhigh reasoning
59.10%
70
Claude Opus 4Anthropic
58.59%
71
58.51%
76
GPT-4oOpenAI · Closedhigh reasoning
57.43%
77
56.76%
80
GPT-4o miniOpenAI · Closedhigh reasoning
54.49%
81
GPT-5 nanoOpenAI · Closedhigh reasoning
53.62%
82
GPT-4.1 nanoOpenAI · Closedhigh reasoning
52.82%
83
52.11%
84
51.87%
86
Grok 4.3xAI · Closed
48.25%
88
44.48%
89
42.77%
90
42.61%
91
42.09%
93
36.45%
95
Mistral Medium 3.5Mistral AIhigh reasoning
28.89%

How MortgageTax is shown here

BenchLM mirrors the public Vals AI MortgageTax leaderboard captured from https://www.vals.ai/benchmarks/mortgage_tax and updated by Vals on September 1, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

MortgageTax is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

98 Vals rows3 task viewsprivate datasetTasks: Overall, Semantic Extraction, Numerical ExtractionDisplay only

The published MortgageTax snapshot places Claude Opus 5 first at 72.06%. The third row is 1.79 points behind. The broader top-10 range is 3.30 points, so many of the published results sit in a relatively narrow band.

98 models have been evaluated on MortgageTax. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. MortgageTax is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MortgageTax

Year

2026

Tasks

Mortgage and tax extraction tasks

Format

Accuracy score

Difficulty

Professional mortgage-tax document reasoning

BenchLM mirrors Vals MortgageTax as a display-only finance and document-reasoning benchmark.

Freshness and provenance

Version

MortgageTax 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MortgageTax measure?

Vals AI benchmark for mortgage and tax document reasoning, including semantic and numerical extraction task views.

Which model leads the published MortgageTax snapshot?

Claude Opus 5 currently leads the published MortgageTax snapshot with 72.06% mortgagetax score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on MortgageTax?

The September 1, 2026 snapshot contains 98 AI models.

Last updated: September 1, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.