Skip to main content
BenchLM

PerceptionBench (Internal) (PerceptionBench)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Moonshot AI's internal benchmark for atomic visual perception capabilities.

Benchmark score on PerceptionBench — September 27, 2026

We compile the PerceptionBench rows from provider self-reports. Qwen3.8 Max leads the table at 63.5%, followed by Kimi K3 (58.5%) and dots3-note Preview (53.4%). We do not use these results to rank models overall.

3 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (3 models)

Score
1
Qwen3.8 MaxAlibaba · Open weight
63.5%
2
Kimi K3Moonshot AI · Closed
58.5%
3
dots3-note PreviewDots Studio · Open weight
53.4%

Among the reported PerceptionBench rows, Qwen3.8 Max is first at 63.5%. The third row is 10.1 points behind. The broader top-10 range is 10.1 points, so the table still separates the published systems.

3 models have been evaluated on PerceptionBench. The benchmark falls in the Multimodal & Grounded category. PerceptionBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About PerceptionBench

Year

2026

Tasks

Internal atomic visual-perception tasks

Format

Internal evaluation score

Difficulty

Fine-grained visual perception

BenchLM stores the provider-published internal benchmark value as display-only evidence and does not treat it as independently reproducible.

Freshness and provenance

Version

PerceptionBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does PerceptionBench measure?

Moonshot AI's internal benchmark for atomic visual perception capabilities.

Which model scores highest on PerceptionBench?

Qwen3.8 Max by Alibaba currently leads with a score of 63.5% on PerceptionBench.

How many models are evaluated on PerceptionBench?

3 AI models have been evaluated on PerceptionBench on BenchLM.

Compare top models on PerceptionBench

Last updated: September 27, 2026 · BenchLM version PerceptionBench 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.