Skip to main content
BenchLM
Data

Vals Index v2 (Vals Index)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Vals AI composite benchmark across professional finance, coding, modeling, and legal-work tasks, including Finance Agent v2, EMB, Terminal-Bench 2.1, Vibe Code Bench, Code Migration, Legal Research, and HLAB.

Evaluation results published by Vals AI. Published leaderboard aggregates and evaluation settings only. Dataset contents are not included.

Vals Index score on Vals Index — October 6, 2026

We mirror the published vals index score view for Vals Index. Gemini 4 Argon leads the public snapshot at 68.90%, followed by Claude Sonnet 5.5 (67.04%) and Claude Opus 5.5 (66.97%). We do not use these results to rank models overall.

44 modelsCross-domainCurrentDisplay onlyUpdated October 6, 2026

Vals Index score table (44 models)

Vals Index score results for Vals Index
RankModel / configurationScoreParameters (B)Open / closed
1Gemini 4 ArgonGooglehigh reasoning
68.90%
Not reportedClosed
2Claude Sonnet 5.5Anthropic
67.04%
Not reportedClosed
3Claude Opus 5.5Anthropic
66.97%
Not reportedClosed
4Claude Fable 5.1Anthropic
65.83%
Not reportedClosed
5Claude Opus 5Anthropic
63.67%
Not reportedClosed
6GPT-6 AstraOpenAImax reasoning
63.13%
Not reportedClosed
7Claude Fable 5Anthropic
61.39%
Not reportedClosed
8GPT-6.1 SolOpenAImax reasoning
61.15%
Not reportedClosed
9Muse Spark 1.3 MaxMetamax reasoning
58.16%
Not reportedUnknown
10GPT-5.6 SolOpenAImax reasoning
58.01%
Not reportedClosed
11GPT-6 SolOpenAImax reasoning
57.54%
Not reportedClosed
12MiMo-V2.6-ProXiaomi
55.20%
Not reportedOpen
13Claude Opus 4.8Anthropic
55.10%
Not reportedClosed
14Grok 4.7xAIxhigh reasoning
54.95%
Not reportedClosed
15Gemini 3.8 FlashGooglehigh reasoning
54.83%
Not reportedClosed
16GLM-5.3Z.AImax reasoning
53.51%
Not reportedOpen
17MiMo-V2.6-FlashXiaomi
53.23%
Not reportedOpen
18Muse Spark 1.3Metaxhigh reasoning
53.20%
Not reportedClosed
19GPT-5.6 TerraOpenAImax reasoning
53.09%
Not reportedClosed
20Grok 4.6xAIhigh reasoning
52.09%
Not reportedClosed
21Claude Sonnet 5Anthropic
51.77%
Not reportedClosed
22GPT-5.6 LunaOpenAImax reasoning
51.69%
Not reportedClosed
23DeepSeek V4.1 FlashDeepSeekhigh reasoning
51.32%
Not reportedOpen
24Gemini 3.7 FlashGooglehigh reasoning
51.27%
Not reportedClosed
25GPT-6 LunaOpenAImax reasoning
51.22%
Not reportedClosed
26Ember-1Fireworks
50.81%
Not reportedClosed
27Kimi K3Moonshot AImax reasoning
50.30%
Not reportedPending
28Hy4 previewTencent
49.94%
Not reportedOpen
29Step 5 PreviewStepFun
49.35%
Not reportedPending
30Muse Spark 1.2Metaxhigh reasoning
49.29%
Not reportedClosed
31Qwen3.8 MaxAlibaba
48.27%
Not reportedOpen
32Mistral Large 4Mistralhigh reasoning
48.05%
Not reportedPending
33DeepSeek V4 Flash 0731DeepSeekhigh reasoning
47.96%
Not reportedOpen
34DeepSeek V4 Pro 0813DeepSeekmax reasoning
47.63%
Not reportedOpen
35Gemini 3.5 FlashGooglehigh reasoning
44.79%
Not reportedClosed
36Grok 4.5xAIhigh reasoning
44.72%
Not reportedClosed
37DeepSeek V4 Pro 0813DeepSeekmax reasoning
38.63%
Not reportedOpen
38MiniMax M3MiniMax
36.53%
Not reportedOpen
39MiMo-V2.5-ProXiaomi
33.84%
Not reportedClosed
40Gemini 3.1 Pro PreviewGooglehigh reasoning
33.44%
Not reportedUnknown
41GPT-5.4 miniOpenAIxhigh reasoning
33.17%
Not reportedClosed
42InklingThinking Machines Lab0.99 reasoning
28.67%
Not reportedOpen
43Inkling-SmallThinking Machines Lab0.99 reasoning
25.46%
Not reportedOpen
44Mercury 2.5Inceptionhigh reasoning
8.86%
Not reportedClosed

How Vals Index is shown here

BenchLM mirrors the public Vals AI Vals Index leaderboard captured from https://www.vals.ai/benchmarks/vals_index and updated by Vals on October 6, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Index is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

44 Vals rows9 task viewsprivate datasetTasks: Overall, Finance Agent v2, EMB, Terminal-Bench 4.0, Vibe Code BenchDisplay only

The published Vals Index snapshot places Gemini 4 Argon first at 68.90%. The third row is 1.92 points behind. The broader top-10 range is 10.89 points, so the table still separates the published systems.

44 models have been evaluated on Vals Index. The benchmark falls in the Cross-domain category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals Index is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vals Index

Year

2026

Tasks

Finance, coding, spreadsheet modeling, code migration, and legal-work components

Format

Composite score

Difficulty

Private economic-work benchmark composite

We mirror Vals Index v2 as a display-only external composite. The task mix changed from the earlier index, so the current score should not be treated as directly interchangeable with an older snapshot. Vals proprietary aggregate scores remain outside weighted rankings.

Freshness and provenance

Version

Vals Index 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Vals Index measure?

Vals AI composite benchmark across professional finance, coding, modeling, and legal-work tasks, including Finance Agent v2, EMB, Terminal-Bench 2.1, Vibe Code Bench, Code Migration, Legal Research, and HLAB.

Which model leads the published Vals Index snapshot?

Gemini 4 Argon currently leads the published Vals Index snapshot with 68.90% vals index score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vals Index?

The October 6, 2026 snapshot contains 44 AI models.

Last updated: October 6, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.