Skip to main content
BenchLM
Data

Finance Agent v2

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Gemini 4 Argon leads Vals' October 6, 2026 Finance Agent v2 table at 65.4% partial credit. Its all-pass score is 55.8%. The test checks financial research with tools; partial credit rewards weighted checks after required dealbreakers pass, while all-pass requires every check.

Evaluation results published by Vals AI. Published leaderboard aggregates and evaluation settings only. Dataset contents are not included.

Finance Agent v2 score on Finance Agent v2 — October 6, 2026

We mirror the published finance agent v2 score view for Finance Agent v2. Gemini 4 Argon leads the public snapshot at 65.4%, followed by Gemini 3.8 Flash (61.4%) and Muse Spark 1.2 (60.6%). We do not use these results to rank models overall.

75 modelsAgenticCurrentDisplay onlyUpdated October 6, 2026

Finance Agent v2 score table (75 models)

Finance Agent v2 score results for Finance Agent v2
RankModel / configurationScoreParameters (B)Open / closed
1Gemini 4 ArgonGooglehigh reasoning
65.4%
Not reportedClosed
2Gemini 3.8 FlashGooglehigh reasoning
61.4%
Not reportedClosed
3Muse Spark 1.2Metaxhigh reasoning
60.6%
Not reportedClosed
4Muse Spark 1.3 MaxMetamax reasoning
60.0%
Not reportedUnknown
5Gemini 3.7 FlashGooglehigh reasoning
59.0%
Not reportedClosed
6Muse Spark 1.3Metaxhigh reasoning
58.9%
Not reportedClosed
7Claude Fable 5.1Anthropic
58.9%
Not reportedClosed
8Claude Opus 5Anthropic
58.6%
Not reportedClosed
9Claude Opus 5.5Anthropic
58.6%
Not reportedClosed
10Claude Sonnet 5.5Anthropic
58.1%
Not reportedClosed
11Gemini 3.5 FlashGooglehigh reasoning
57.9%
Not reportedClosed
12GLM-5.3-FlashZ.AImax reasoning
57.9%
Not reportedOpen
13MiMo-V2.6-ProXiaomi
57.3%
Not reportedOpen
14Muse Spark 1.1Metaxhigh reasoning
57.2%
Not reportedClosed
15Claude Fable 5Anthropic
56.3%
Not reportedClosed
16Gemini 3.6 FlashGooglehigh reasoning
56.3%
Not reportedClosed
17MiMo-V2.6-FlashXiaomi
56.3%
Not reportedOpen
18GLM-5.3Z.AImax reasoning
55.8%
Not reportedOpen
19Hy4 previewTencent
55.1%
Not reportedOpen
20GPT-5.6 LunaOpenAImax reasoning
55.0%
Not reportedClosed
21Ling 3.0 Flash Af Rc3Ant
54.9%
Not reportedUnknown
22Mistral Large 4Mistralhigh reasoning
54.7%
Not reportedPending
23GPT-5.6 TerraOpenAImax reasoning
54.4%
Not reportedClosed
24Claude Opus 4.8Anthropic
53.9%
Not reportedClosed
25Claude Sonnet 5Anthropic
53.9%
Not reportedClosed
26GPT-5.6 SolOpenAImax reasoning
53.8%
Not reportedClosed
27Grok 4.6xAIhigh reasoning
53.7%
Not reportedClosed
28GPT-6 AstraOpenAImax reasoning
53.5%
Not reportedClosed
29DeepSeek V4.1 FlashDeepSeekhigh reasoning
53.5%
Not reportedOpen
30Kimi K3Moonshot AImax reasoning
53.1%
Not reportedPending
31Grok 4.7xAIxhigh reasoning
52.3%
Not reportedClosed
32GPT-6.1 SolOpenAImax reasoning
52.0%
Not reportedClosed
33Ember-1Fireworks
51.9%
Not reportedClosed
34GPT-5.5OpenAIxhigh reasoning
51.8%
Not reportedClosed
35Claude Opus 4.7Anthropic
51.5%
Not reportedClosed
36Claude Sonnet 4.6Anthropic
51.0%
Not reportedClosed
37Step 5 PreviewStepFun
50.7%
Not reportedPending
38Qwen3.8 MaxAlibaba
50.6%
Not reportedOpen
39DeepSeek V4 Pro 0813DeepSeekmax reasoning
50.4%
Not reportedOpen
40GPT-6 LunaOpenAImax reasoning
49.9%
Not reportedClosed
41GLM-5.2Z.AI
49.7%
Not reportedOpen
42DeepSeek V4 Flash 0731DeepSeekhigh reasoning
49.5%
Not reportedOpen
43GPT-6 SolOpenAImax reasoning
49.0%
Not reportedClosed
44Qwen3.8-27BAlibabaxhigh reasoning
48.6%
Not reportedOpen
45Grok 4.5xAIhigh reasoning
48.3%
Not reportedClosed
46MiniMax M3MiniMax
48.3%
Not reportedOpen
47Qwen3.7 MaxAlibaba
47.8%
Not reportedClosed
48Gemini 3.5 Flash-LiteGooglehigh reasoning
47.4%
Not reportedClosed
49InklingThinking Machines Lab0.99 reasoning
46.6%
Not reportedOpen
50GPT-5.4 miniOpenAIxhigh reasoning
45.4%
Not reportedClosed
51Kimi K2.6Moonshot AI
44.9%
Not reportedOpen
52GLM-5.1Z.AI
44.8%
Not reportedOpen
53DeepSeek V4 Pro 0813DeepSeekmax reasoning
44.1%
Not reportedOpen
54Gemini 3.1 Pro PreviewGooglehigh reasoning
43.0%
Not reportedUnknown
55Gemini 3 Flash PreviewGooglehigh reasoning
42.6%
Not reportedUnknown
56MiMo-V2.5-ProXiaomi
41.5%
Not reportedClosed
57Inkling-SmallThinking Machines Lab0.99 reasoning
41.3%
Not reportedOpen
58Qwen3.6 PlusAlibaba
40.8%
Not reportedClosed
59Qwen3.7 PlusAlibaba
38.2%
Not reportedClosed
60GPT-5.4 nanoOpenAIhigh reasoning
38.2%
Not reportedClosed
61Grok 4.3xAIhigh reasoning
37.7%
Not reportedClosed
62Nemotron 3 Ultra 550b A55bNvidia
37.7%
Not reportedUnknown
63MiMo-V2.5Xiaomi
36.7%
Not reportedClosed
64Kimi K2.5 ThinkingMoonshot AI
35.8%
Not reportedUnknown
65Mistral Medium 3.5Mistral AIhigh reasoning
32.1%
Not reportedUnknown
66Claude Haiku 4.5 ThinkingAnthropic
31.0%
Not reportedClosed
67Ling 3.0 Flash 2607Ant
30.3%
Not reportedUnknown
68Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
30.0%
Not reportedUnknown
69Grok 4.20 0309 ReasoningxAI
28.5%
Not reportedUnknown
70MiniMax M2.7MiniMax
27.9%
Not reportedOpen
71Laguna M.1Poolside
25.0%
Not reportedClosed
72Mercury 2.5Inceptionhigh reasoning
18.9%
Not reportedClosed
73Nemotron Lightning 3p5 30b A3bFireworks AI
18.5%
Not reportedUnknown
74Laguna XS.2Poolside
15.6%
Not reportedOpen
75Command A Plus 05 2026Cohere
9.0%
Not reportedUnknown

How Finance Agent v2 is shown here

BenchLM mirrors the public Vals AI Finance Agent v2 leaderboard captured from https://www.vals.ai/benchmarks/fabv2 and updated by Vals on October 6, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Finance Agent v2 is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

75 Vals rows11 task viewsprivate datasetTasks: Overall, All-Pass, General Qualitative Analysis, General Quantitative Analysis, Market AnalysisDisplay only

The published Finance Agent v2 snapshot places Gemini 4 Argon first at 65.4%. The third row is 4.8 points behind. The broader top-10 range is 7.3 points, so many of the published results sit in a relatively narrow band.

75 models have been evaluated on Finance Agent v2. The benchmark falls in the Agentic category. Finance Agent v2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Finance Agent v2

Year

2026

Tasks

Financial analyst task categories

Format

Mean score across repeated runs

Difficulty

Professional expert-task agent workflow

Vals reports Finance Agent v2 as a multi-category benchmark with severity-weighted partial credit and repeated runs per model. BenchLM mirrors the public Vals leaderboard as a display-only expert-task benchmark.

Freshness and provenance

Version

Finance Agent v2 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Finance Agent v2 measure?

Vals AI benchmark for realistic financial analyst agent tasks across qualitative analysis, quantitative analysis, market work, comparables, precedents, earnings, disclosure, and modeling.

Which model leads the published Finance Agent v2 snapshot?

Gemini 4 Argon currently leads the published Finance Agent v2 snapshot with 65.4% finance agent v2 score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Finance Agent v2?

The October 6, 2026 snapshot contains 75 AI models.

Last updated: October 6, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.