Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

ScreenSpot Pro

A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.

Data verified 24 confirmed releases in the last 30 daysStart free brief

The public ScreenSpot Pro snapshot ranks Claude Opus 4.8 first at 87.9%, ahead of GPT-5.4 (85.4%) and Qwen3.8 Max (84.5%) among 17 models. We mirror the table as display-only evidence; it does not affect overall rankings.

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Use ScreenSpot Pro to judge published GUI-grounding systems after checking how each system handles the image. Direct coordinate prediction, cropping, visual search, a planner, or Python tools can change the result without changing the underlying base model.

Operator receipt: 17 sourced rows are currently displayable on this page; the leading published row is Claude Opus 4.8 at 87.9%.

Honest limit: ScreenSpot Pro stops at localization on a static screenshot. It does not test clicking, typing, state tracking, recovery, or completion of a multi-step workflow, and published rows with different tool or search setups are not a clean model-only comparison.

Benchmark score on ScreenSpot Pro — August 26, 2026

We mirror the published score view for ScreenSpot Pro. Claude Opus 4.8 leads the public snapshot at 87.9%, followed by GPT-5.4 (85.4%) and Qwen3.8 Max (84.5%). We do not use these results to rank models overall.

17 modelsMultimodal & GroundedCurrentDisplay onlyUpdated August 26, 2026

Benchmark score table (17 models)

Score
1
Claude Opus 4.8Anthropic · Closed
87.9%
2
GPT-5.4OpenAI · Closed
85.4%
3
Qwen3.8 MaxAlibaba · Open weight
84.5%
4
Gemini 3.1 ProGoogle · Closed
84.4%
5
Muse SparkMeta · Closed
84.1%
6
Claude Opus 4.6Anthropic · Closed
83.1%
7
Qwen3.7 PlusAlibaba · Closed
79.0%
8
Muse Glimmer 30BMeta · Open weight
75.4%
9
Gemini 3 ProGoogle · Closed
72.7%
10
Holo2-235B-A22BH Company · Open weight
70.6%
11
Qwen3.6 PlusAlibaba · Closed
68.2%
12
Holo2-30B-A3BH Company · Open weight
66.1%
13
Qwen3.5 397BAlibaba · Open weight
65.6%
14
Holo2-8BH Company · Open weight
58.9%
15
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
57.8%
16
Holo2-4BH Company · Open weight
57.2%
17
Claude Opus 4.5Anthropic · Closed
45.7%

The published ScreenSpot Pro snapshot places Claude Opus 4.8 first at 87.9%. The third row is 3.4 points behind. The broader top-10 range is 17.3 points, so the table still separates the published systems.

17 models have been evaluated on ScreenSpot Pro. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. ScreenSpot Pro is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ScreenSpot Pro

Year

2025

Tasks

1,581 grounding instructions

Format

Static interface element localization

Difficulty

Professional GUI grounding

The benchmark spans 23 applications, six application categories, and three operating systems. It scores whether a system localizes the requested text or icon in a static professional screenshot, using micro-average accuracy on the official leaderboard.

BenchLM freshness & provenance

Version

ScreenSpot Pro 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ScreenSpot Pro measure?

ScreenSpot Pro measures whether a model can map a natural-language instruction to the correct target in a full-screen, high-resolution professional interface. Its 1,581 examples span 23 applications, six application categories, and three operating systems. The score is grounding accuracy, not end-to-end task completion.

What does a high ScreenSpot Pro score mean?

A high score means the published system localized more requested interface elements under its reported setup. It does not prove the same base model will click correctly inside an agent. Cropping, visual search, planners, Python tools, image resolution, and decoding policy can materially change the result.

Can ScreenSpot Pro pick the best computer-use agent?

No. ScreenSpot Pro tests static-screen grounding, which is one prerequisite for computer use. It does not test state tracking, typing, recovery after a bad action, or completion of a multi-step workflow. Pair it with OSWorld Verified and an application-specific operator test before choosing an agent.

Last updated: August 26, 2026 · BenchLM version ScreenSpot Pro 2025

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.