Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Vals Finance Agent v1.1 (Finance Agent v1.1)

An archived Vals benchmark covering retrieval, numerical reasoning, financial modeling, market analysis, earnings, trends, and adjustments.

Data verified 26 confirmed releases in the last 30 daysStart free brief

How BenchLM shows Finance Agent v1.1

BenchLM mirrors the public Vals AI Finance Agent v1.1 leaderboard captured from https://www.vals.ai/benchmarks/finance_agent and updated by Vals on June 4, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Finance Agent v1.1 is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

51 Vals rows11 task viewsprivate datasetTasks: Overall, Simple retrieval - Quantitative, Simple retrieval - Qualitative, Complex Retrieval, Numerical ReasoningDisplay only

Finance Agent v1.1 score on Finance Agent v1.1 — June 4, 2026

We mirror the published finance agent v1.1 score view for Finance Agent v1.1. Claude Opus 4.7 leads the public snapshot at 64.37%, followed by Claude Sonnet 4.6 (63.33%) and Muse Spark (60.59%). We do not use these results to rank models overall.

51 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated June 4, 2026

Finance Agent v1.1 score table (51 models)

Score
1
Claude Opus 4.7Anthropic · Closed
64.37%
2
Claude Sonnet 4.6Anthropic · Closed
63.33%
3
Muse SparkMeta · Closed
60.59%
4
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
60.39%
5
Claude Opus 4.6 (Adaptive)Anthropic · Closed
60.05%
6
GPT-5.5OpenAI · Closedxhigh reasoning
59.96%
7
Gemini 3.1 Pro PreviewGooglehigh reasoning
59.72%
8
Claude Opus 4.5 ThinkingAnthropic · Closed
58.81%
9
GPT-5.2OpenAI · Closedxhigh reasoning
58.53%
10
GLM-5.1Z.AI · Open weight
57.66%
11
GPT-5.4OpenAI · Closedxhigh reasoning
57.15%
12
Kimi K2.6Moonshot AI · Open weight
57.06%
13
GPT-5.1OpenAI · Closedhigh reasoning
55.31%
14
Gemini 3 Pro PreviewGooglehigh reasoning
55.15%
15
Qwen3.6 PlusAlibaba · Closed
54.63%
16
Claude Sonnet 4.5 ThinkingAnthropic · Closed
54.50%
17
54.48%
18
Grok 4.3xAI · Closed
53.81%
19
53.51%
20
GPT-5.4 miniOpenAI · Closedxhigh reasoning
53.41%
21
53.18%
22
Qwen 3.6 Max (preview)Alibaba · Closed
52.78%
23
52.45%
25
GPT-5OpenAIhigh reasoning
52.15%
26
GPT-5 miniOpenAI · Closedhigh reasoning
51.93%
27
Gemma 4 31b ItGooglehigh reasoning
50.79%
28
50.62%
29
MiniMax M2.7MiniMax · Open weight
48.40%
30
GPT-5.4 nanoOpenAI · Closedhigh reasoning
47.80%
31
Gemini 3 Flash PreviewGooglehigh reasoning
47.60%
32
Claude Haiku 4.5 ThinkingAnthropic · Closed
46.93%
33
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
46.12%
34
Mistral Medium 3.5Mistral AIhigh reasoning
46.11%
35
46.08%
36
GLM-4.7Z.AI · Open weight
45.98%
37
Qwen3.5 FlashAlibaba · Closed
45.64%
39
Qwen3 MaxAlibaba · Closed
44.30%
40
Gemini 2.5 ProGoogle · Closed
41.59%
41
MiniMax M2.5MiniMax · Closed
38.58%
42
Kimi K2 ThinkingMoonshot AI
36.65%
43
GLM-4.6Z.AI · Open weight
36.48%
44
33.35%
45
GPT-OSS 120BOpenAI · Open weight
21.54%
46
18.05%
47
GPT-4oOpenAI · Closedhigh reasoning
8.06%
48
4.23%
49
DeepSeek V3p2 ThinkingFireworks AIhigh reasoning
2.35%
50
0.37%
51
DeepSeek V3p2Fireworks AInone reasoning
0.00%

The published Finance Agent v1.1 snapshot places Claude Opus 4.7 first at 64.37%. The third row is 3.78 points behind. The broader top-10 range is 6.72 points, so many of the published results sit in a relatively narrow band.

51 models have been evaluated on Finance Agent v1.1. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Finance Agent v1.1 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Finance Agent v1.1

Year

2026

Tasks

11 financial analyst task views

Format

Accuracy with task-level breakdowns

Difficulty

Professional financial analysis

We mirror all 51 published rows and 11 task views for historical comparison. Finance Agent v2 is the current Vals benchmark, so v1.1 stays display only and does not enter weighted rankings.

BenchLM freshness & provenance

Version

Finance Agent v1.1 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Finance Agent v1.1 measure?

An archived Vals benchmark covering retrieval, numerical reasoning, financial modeling, market analysis, earnings, trends, and adjustments.

Which model leads the published Finance Agent v1.1 snapshot?

Claude Opus 4.7 currently leads the published Finance Agent v1.1 snapshot with 64.37% finance agent v1.1 score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Finance Agent v1.1?

51 AI models are included in BenchLM's mirrored Finance Agent v1.1 snapshot, based on the public leaderboard captured on June 4, 2026.

Last updated: June 4, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.