Skip to main content
BenchLM

Scale Labs Visual-Language Understanding (Visual-Language Understanding)

We show this table for reference; we do not rank on it.

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Scale score on Visual-Language Understanding — September 29, 2026 snapshot

We mirror the published scale score view for Visual-Language Understanding. Gemini 2.5 Pro Experimental (March 2025) leads the public snapshot at 54.6%, followed by gemini-2.5-pro-preview-06-05 (54.6%) and GPT-5.4 Pro (53.9%). We do not use these results to rank models overall.

63 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 29, 2026 snapshot

Scale score table (63 models)

Score
3
GPT-5.4 ProOpenAI · Closed
53.9%
11
GPT-5 miniOpenAI · Closed
50.4%
13
49.7%
16
Claude Sonnet 4.5 ThinkingAnthropic · Closed
48.8%
19
47.3%
21
47.0%
23
GPT-5.2OpenAI · Closed
46.6%
24
Claude Opus 4.5 ThinkingAnthropic · Closed
46.4%
29
GPT-4.1OpenAI · Closed
45.3%
30
Claude Opus 4.5Anthropic · Closed
45.3%
31
45.3%
32
45.3%
33
Claude Sonnet 4.5Anthropic · Closed
45%
34
43.8%
35
Claude Opus 4Anthropic
43.5%
37
43.2%
40
Kimi K2.5Moonshot AI · Open weight
41.9%
41
GPT-4.1 miniOpenAI · Closed
41.1%
46
Llama 4 MaverickMeta · Open weight
38.3%
48
Gemini 1.5 ProGoogle · Closed
37.1%
50
34.9%
51
Mistral Medium 3Mistral · Closed
34.6%
55
28.6%
56
Claude 3 OpusAnthropic · Closed
27.8%
57
GPT-4.1 nanoOpenAI · Closed
26.6%
58
Nova ProAmazon · Closed
26.3%
60
Nova LiteAmazon
25.5%
63
15.2%

How to read this leaderboard

Compare the published configurations as complete evaluation systems. The source can combine a base model, agent scaffold, tools, budget, and inference setting in each result.

Operator receipt: 63 sourced rows are currently displayable on this page; the leading published row is Gemini 2.5 Pro Experimental (March 2025) at 54.6%.

Honest limit: This Scale table is display-only context, not benchmark provenance or a weighted model-only comparison.

How BenchLM shows Visual-Language Understanding

BenchLM mirrors 63 published rows from Scale Labs’ public Visual-Language Understanding leaderboard, captured on September 29, 2026 snapshot.

The table is display only. It is useful context for a published agent or model configuration, but it does not enter BenchLM’s overall or category rankings.

Snapshot

63 published rowsScale Labs sourceDisplay only

The published Visual-Language Understanding snapshot places Gemini 2.5 Pro Experimental (March 2025) first at 54.6%. The third row is 0.8 points behind. The broader top-10 range is 3.9 points, so many of the published results sit in a relatively narrow band.

63 models have been evaluated on Visual-Language Understanding. The benchmark falls in the Multimodal & Grounded category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Visual-Language Understanding is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Visual-Language Understanding

Year

2026

Tasks

63 published rows

Format

Published Scale leaderboard score

Difficulty

External agent and model evaluation

BenchLM mirrors 63 published rows from the Visual-Language Understanding public table captured on September 29, 2026 snapshot.

Freshness and provenance

Version

Visual-Language Understanding 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Visual-Language Understanding measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Which model leads the published Visual-Language Understanding snapshot?

Gemini 2.5 Pro Experimental (March 2025) currently leads the published Visual-Language Understanding snapshot with 54.6% scale score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Visual-Language Understanding?

The September 29, 2026 snapshot snapshot contains 63 AI models.

Last updated: September 29, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.