Skip to main content
BenchLM

V*

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A vision-centric benchmark for high-level multimodal reasoning and perception quality.

Benchmark score on V* — September 27, 2026

We compile the V* rows from provider self-reports and secondary reports. Kimi K2.6 leads the table at 96.9%, followed by Qwen3.6 Plus (96.9%) and Qwen3.5 397B (95.8%). We do not use these results to rank models overall.

11 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (11 models)

Score
1
Kimi K2.6Moonshot AI · Open weight
96.9%
2
Qwen3.6 PlusAlibaba · Closed
96.9%
3
Qwen3.5 397BAlibaba · Open weight
95.8%
4
Step 3.7 FlashStepFun · Open weight
95.3%
5
Qwen3.6-27BAlibaba · Open weight
94.7%
6
Qwen3.5-27BAlibaba · Open weight
93.7%
7
Qwen3.5-122B-A10BAlibaba · Open weight
93.2%
8
Qwen3.5-35B-A3BAlibaba · Open weight
92.7%
9
Gemini 3 ProGoogle · Closed
88.0%
10
GPT-5.2OpenAI · Closed
75.9%
11
Claude Opus 4.5Anthropic · Closed
67.0%

Among the reported V* rows, Kimi K2.6 is first at 96.9%. The third row is 1.1 points behind. The broader top-10 range is 21.0 points, so the table still separates the published systems.

11 models have been evaluated on V*. The benchmark falls in the Multimodal & Grounded category. V* is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About V*

Year

2026

Tasks

Frontier multimodal reasoning tasks

Format

Vision-centric reasoning benchmark

Difficulty

Frontier multimodal

BenchLM tracks V* as a display-only frontier multimodal benchmark reference outside the current weighted schema.

Freshness and provenance

Version

V* 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does V* measure?

A vision-centric benchmark for high-level multimodal reasoning and perception quality.

Which model scores highest on V*?

Kimi K2.6 by Moonshot AI currently leads with a score of 96.9% on V*.

How many models are evaluated on V*?

11 AI models have been evaluated on V* on BenchLM.

Last updated: September 27, 2026 · BenchLM version V* 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.