# Vals Index v2 (Vals Index)

> Vals AI composite benchmark across professional finance, coding, modeling, and legal-work tasks, including Finance Agent v2, EMB, Terminal-Bench 2.1, Vibe Code Bench, Code Migration, Legal Research, and HLAB.

Canonical page: https://benchlm.ai/benchmarks/valsindex

- Category: [external](/external)
- Last updated: September 15, 2026

## About Vals Index

- Year: 2026
- Tasks: Finance, coding, spreadsheet modeling, code migration, and legal-work components
- Format: Composite score
- Difficulty: Private economic-work benchmark composite
- Paper: [Vals Index](https://www.vals.ai/benchmarks/vals_index)

We mirror Vals Index v2 as a display-only external composite. The task mix changed from the earlier index, so the current score should not be treated as directly interchangeable with an older snapshot. Vals proprietary aggregate scores remain outside weighted rankings.

Vals Index is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (58 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | — | Anthropic | 68.83% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | — | Anthropic | 67.21% |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | max reasoning | OpenAI | 66.61% |
| 4 | [Claude Fable 5](/models/claude-fable) | — | Anthropic | 66.04% |
| 5 | [Muse Spark 1.3 Max](https://www.vals.ai/models/meta_muse_spark_1_3_max) | max reasoning | Meta | 64.53% |
| 6 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | max reasoning | OpenAI | 63.71% |
| 7 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | high reasoning | Google | 62.25% |
| 8 | [Claude Opus 4.8](/models/claude-opus-4-8) | — | Anthropic | 60.91% |
| 9 | [Muse Spark 1.3](/models/muse-spark-1-3) | xhigh reasoning | Meta | 60.31% |
| 10 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | max reasoning | OpenAI | 59.88% |
| 11 | [Claude Sonnet 5](/models/claude-sonnet-5) | — | Anthropic | 59.61% |
| 12 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | max reasoning | OpenAI | 59.59% |
| 13 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | high reasoning | Google | 59.31% |
| 14 | [Grok 4.6](/models/grok-4-6) | high reasoning | xAI | 59.17% |
| 15 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | high reasoning | DeepSeek | 57.86% |
| 16 | [Kimi K3](/models/kimi-k3) | — | Moonshot AI | 57.81% |
| 17 | [GPT-5.5](/models/gpt-5-5) | xhigh reasoning | OpenAI | 57.41% |
| 18 | [Muse Spark 1.2](/models/muse-spark-1-2) | xhigh reasoning | Meta | 57.05% |
| 19 | [GLM-5.3](/models/glm-5-3) | max reasoning | Z.AI | 56.97% |
| 20 | [Claude Opus 4.7](/models/claude-opus-4-7) | — | Anthropic | 56.11% |
| 21 | [Hy4 preview](/models/hy4-preview) | — | Tencent | 55.40% |
| 22 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | high reasoning | Google | 55.35% |
| 23 | [Muse Spark 1.1](/models/muse-spark-1-1) | xhigh reasoning | Meta | 54.75% |
| 24 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | high reasoning | DeepSeek | 53.57% |
| 25 | [GLM-5.2](/models/glm-5-2) | — | Z.AI | 53.12% |
| 26 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | high reasoning | Google | 53.08% |
| 27 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | max reasoning | DeepSeek | 52.37% |
| 28 | [Qwen3.8 Max](/models/qwen3-8-max) | — | Alibaba | 51.84% |
| 29 | [Grok 4.5](/models/grok-4-5) | high reasoning | xAI | 51.53% |
| 30 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | — | Anthropic | 50.59% |
| 31 | [Qwen3.8-27B](/models/qwen3-8-27b) | xhigh reasoning | Alibaba | 48.48% |
| 32 | [GLM-5.3-Flash](/models/glm-5-3-flash) | max reasoning | Z.AI | 47.22% |
| 33 | [Qwen3.7 Max](/models/qwen3-7-max) | — | Alibaba | 44.77% |
| 34 | [Kimi K2.6](/models/kimi-2-6) | — | Moonshot AI | 43.47% |
| 35 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | max reasoning | DeepSeek | 42.89% |
| 36 | [MiniMax M3](/models/minimax-m3) | — | MiniMax | 42.72% |
| 37 | [Gemini 3.1 Pro Preview](https://www.vals.ai/models/google_gemini-3.1-pro-preview) | high reasoning | Google | 41.90% |
| 38 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | — | Xiaomi | 40.97% |
| 39 | [MiMo-V2.5](/models/mimo-v2-5) | — | Xiaomi | 39.91% |
| 40 | [GPT-5.4 mini](/models/gpt-5-4-mini) | xhigh reasoning | OpenAI | 39.63% |
| 41 | [Qwen3.7 Plus](/models/qwen3-7-plus) | — | Alibaba | 38.65% |
| 42 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | high reasoning | Google | 36.71% |
| 43 | [Inkling](/models/inkling) | 0.99 reasoning | Thinking Machines Lab | 34.10% |
| 44 | [GPT-5.4 nano](/models/gpt-5-4-nano) | high reasoning | OpenAI | 33.00% |
| 45 | [Qwen3.6 Plus](/models/qwen3-6-plus) | — | Alibaba | 31.98% |
| 46 | [Inkling-Small](/models/inkling-small) | 0.99 reasoning | Thinking Machines Lab | 31.86% |
| 47 | [Gemini 3 Flash Preview](https://www.vals.ai/models/google_gemini-3-flash-preview) | high reasoning | Google | 29.44% |
| 48 | [Nemotron 3 Ultra 550b A55b](https://www.vals.ai/models/nvidia_nemotron-3-ultra-550b-a55b) | — | Nvidia | 27.39% |
| 49 | [Kimi K2.5 Thinking](https://www.vals.ai/models/kimi_kimi-k2.5-thinking) | — | Moonshot AI | 26.30% |
| 50 | [MiniMax M2.7](/models/minimax-m2-7) | — | MiniMax | 24.56% |
| 51 | [Grok 4.3](/models/grok-4-3) | high reasoning | xAI | 24.29% |
| 52 | [Claude Haiku 4.5 Thinking](/models/claude-haiku-4-5-thinking) | — | Anthropic | 22.90% |
| 53 | [Ling 3.0 Flash 2607](https://www.vals.ai/models/ant_ling-3.0-flash-2607) | — | Ant | 21.70% |
| 54 | [Mistral Medium 3.5](https://www.vals.ai/models/mistralai_mistral-medium-3.5) | high reasoning | Mistral AI | 17.95% |
| 55 | [Grok 4.20 0309 Reasoning](https://www.vals.ai/models/grok_grok-4.20-0309-reasoning) | — | xAI | 17.55% |
| 56 | [Gemini 3.1 Flash Lite Preview](https://www.vals.ai/models/google_gemini-3.1-flash-lite-preview) | high reasoning | Google | 15.46% |
| 57 | [Mercury 2.5](/models/mercury-2-5) | high reasoning | Inception | 12.95% |
| 58 | [Nemotron Lightning 3p5 30b A3b](https://www.vals.ai/models/fireworks_nemotron-lightning-3p5-30b-a3b) | — | Fireworks AI | 11.49% |

## FAQ

### What does Vals Index measure?

Vals AI composite benchmark across professional finance, coding, modeling, and legal-work tasks, including Finance Agent v2, EMB, Terminal-Bench 2.1, Vibe Code Bench, Code Migration, Legal Research, and HLAB.

### Which model leads the published Vals Index snapshot?

Claude Fable 5.1 currently leads the published Vals Index snapshot with a score of 68.83%.

### How many models are evaluated on Vals Index?

The September 15, 2026 contains 58 AI models.

### Does Vals Index affect BenchLM's overall score?

Not directly. Vals Index is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
