Skip to main content
BenchLM

Vals CorpFin v2 (CorpFin v2)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Vals AI private benchmark for understanding long-context credit agreements.

CorpFin v2 score on CorpFin v2 — August 12, 2026

We mirror the published corpfin v2 score view for CorpFin v2. Claude Opus 5 leads the public snapshot at 73.19%, followed by Claude Fable 5 (71.83%) and Kimi K3 (71.56%). We do not use these results to rank models overall.

134 modelsAgenticCurrentDisplay onlyUpdated August 12, 2026

CorpFin v2 score table (134 models)

Score
1
Claude Opus 5Anthropic · Closed
73.19%
2
Claude Fable 5Anthropic · Closed
71.83%
3
Kimi K3Moonshot AI · Closed
71.56%
4
Muse Spark 1.1Meta · Closedxhigh reasoning
71.29%
5
Muse Spark 1.2Meta · Closedxhigh reasoning
70.94%
6
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
69.62%
7
InklingThinking Machines Lab · Open weight0.99 reasoning
68.57%
8
Grok 4.3xAI · Closed
68.53%
9
GPT-5.5OpenAI · Closedxhigh reasoning
68.42%
10
68.26%
11
MiniMax M3MiniMax · Open weight
68.10%
12
Qwen3 MaxAlibaba · Closed
68.03%
13
Grok 4.5xAI · Closedhigh reasoning
67.41%
14
Claude Opus 4.6 (Adaptive)Anthropic · Closed
67.02%
15
Claude Sonnet 5Anthropic · Closed
66.98%
16
66.90%
17
Kimi K2.6Moonshot AI · Open weight
66.74%
18
Claude Opus 4.8Anthropic · Closed
66.71%
19
Qwen 3.6 Max (preview)Alibaba · Closed
66.47%
20
Gemini 3 Flash PreviewGooglehigh reasoning
66.43%
21
Grok 4.6xAI · Closedhigh reasoning
66.16%
22
GLM-5.2Z.AI · Open weight
66.12%
23
Claude Opus 4.7Anthropic · Closed
66.08%
24
66.05%
25
65.97%
26
GPT-5.2OpenAI · Closedxhigh reasoning
65.89%
27
Qwen3.8 MaxAlibaba · Open weight
65.85%
29
DeepSeek V4 Pro 0813DeepSeek · Open weightmax reasoning
65.42%
30
GPT-5.6 TerraOpenAI · Closedxhigh reasoning
65.31%
31
Claude Sonnet 4.6Anthropic · Closed
65.31%
32
65.31%
33
GPT-5.4OpenAI · Closedxhigh reasoning
65.27%
34
Muse SparkMeta · Closed
65.11%
35
Claude Opus 4.5 ThinkingAnthropic · Closed
65.07%
36
Gemini 3.5 FlashGoogle · Closedhigh reasoning
64.69%
37
Gemini 3.1 Pro PreviewGooglehigh reasoning
64.49%
38
GLM-5.1Z.AI · Open weight
64.45%
39
GPT-5.6 SolOpenAI · Closedmax reasoning
64.38%
40
GPT-5.6 LunaOpenAI · Closedmax reasoning
64.22%
41
GPT-5.1OpenAI · Closedhigh reasoning
63.83%
42
Qwen3.7 MaxAlibaba · Closed
63.71%
44
Gemini 3 Pro PreviewGooglehigh reasoning
63.67%
45
Qwen3.5 FlashAlibaba · Closed
63.56%
46
Gemini 3.6 FlashGoogle · Closedhigh reasoning
63.33%
47
GPT-4.1OpenAI · Closedhigh reasoning
63.05%
48
62.90%
49
Qwen3.6-27BAlibaba · Open weight
62.31%
50
Claude Sonnet 4.5 ThinkingAnthropic · Closed
61.97%
52
Qwen3.6 PlusAlibaba · Closed
61.93%
53
DeepSeek V4 Flash 0731DeepSeek · Open weighthigh reasoning
61.85%
54
MiMo-V2.5-ProXiaomi · Closed
61.42%
55
DeepSeek V4 Pro 0813DeepSeek · Open weightmax reasoning
61.38%
56
Claude Opus 4.5Anthropic · Closed
61.30%
58
GPT-5.4 nanoOpenAI · Closedhigh reasoning
61.19%
59
MiniMax M2.7MiniMax · Open weight
61.19%
60
61.11%
61
GPT-5OpenAIhigh reasoning
61.07%
62
61.03%
63
GLM-4.5Z.AI · Closed
60.96%
64
GPT-5.4 miniOpenAI · Closedxhigh reasoning
60.92%
65
Claude Sonnet 4.5Anthropic · Closed
60.80%
66
Gemini 2.5 ProGoogle · Closed
60.80%
67
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
60.68%
68
Claude Haiku 4.5 ThinkingAnthropic · Closed
60.61%
69
Kimi K2 ThinkingMoonshot AI
60.57%
71
Claude Haiku 4.5Anthropic · Closed
60.30%
72
GPT-5 miniOpenAI · Closedhigh reasoning
60.18%
73
MiMo-V2.5Xiaomi · Closed
59.91%
76
o3OpenAI · Closedhigh reasoning
59.71%
77
59.71%
78
MiniMax M2.5MiniMax · Closed
59.60%
79
59.48%
80
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
59.36%
81
O4 MiniOpenAIhigh reasoning
58.97%
83
58.90%
84
Mistral Medium 3.5Mistral AIhigh reasoning
58.78%
86
GPT-OSS 120BOpenAI · Open weight
58.24%
87
Laguna M.1Poolside · Closed
58.16%
88
GPT-4.1 miniOpenAI · Closedhigh reasoning
57.93%
90
GLM-4.6Z.AI · Open weight
56.84%
91
Laguna XS.2Poolside · Open weight
56.33%
93
Qwen3 MaxAlibaba · Closed
55.94%
94
DeepSeek V3 0324Fireworks AI
54.74%
95
54.70%
96
Trinity-Large-ThinkingArcee AI · Open weight
54.66%
97
54.23%
99
DeepSeek-R1DeepSeek · Open weight
54.12%
100
Claude 3.5 SonnetAnthropic · Closed
53.61%
101
GPT-OSS 20BOpenAI · Open weight
53.15%
102
52.95%
104
DeepSeek V3DeepSeek · Open weight
52.49%
105
DeepSeek V3p1Fireworks AI
51.48%
106
51.13%
107
DeepSeek V3p2 ThinkingFireworks AIhigh reasoning
50.97%
108
50.82%
109
50.74%
110
50.39%
111
49.73%
112
DeepSeek V3p2Fireworks AInone reasoning
47.94%
113
47.40%
115
GLM-4.7Z.AI · Open weight
46.39%
116
45.96%
117
GPT-4oOpenAI · Closedhigh reasoning
45.92%
118
GPT-4o miniOpenAI · Closedhigh reasoning
45.45%
119
o3-miniOpenAI · Closedhigh reasoning
45.30%
120
44.17%
121
44.02%
123
GPT-4.1 nanoOpenAI · Closedhigh reasoning
42.08%
124
41.53%
125
40.52%
126
GPT-4oOpenAI · Closedhigh reasoning
39.43%
127
39.43%
129
38.19%
130
38.03%
132
33.88%
133
33.72%
134
28.63%

How CorpFin v2 is shown here

BenchLM mirrors the public Vals AI CorpFin v2 leaderboard captured from https://www.vals.ai/benchmarks/corp_fin_v2 and updated by Vals on August 12, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

CorpFin v2 is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

134 Vals rows4 task viewsprivate datasetTasks: Overall, Exact Pages, Max Fitting Context, Shared Max ContextDisplay only

The published CorpFin v2 snapshot places Claude Opus 5 first at 73.19%. The third row is 1.63 points behind. The broader top-10 range is 4.94 points, so many of the published results sit in a relatively narrow band.

134 models have been evaluated on CorpFin v2. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. CorpFin v2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About CorpFin v2

Year

2026

Tasks

Credit-agreement understanding tasks

Format

Accuracy score

Difficulty

Professional finance document reasoning

The Vals CorpFin v2 page reports overall, exact-page, max-fitting-context, and shared-max-context task views. BenchLM keeps it display only.

Freshness and provenance

Version

CorpFin v2 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CorpFin v2 measure?

Vals AI private benchmark for understanding long-context credit agreements.

Which model leads the published CorpFin v2 snapshot?

Claude Opus 5 currently leads the published CorpFin v2 snapshot with 73.19% corpfin v2 score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on CorpFin v2?

The August 12, 2026 snapshot contains 134 AI models.

Last updated: August 12, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.