Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

ScreenSpot Pro

A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.

Data verified 37 confirmed releases in the last 30 daysSee provider release alerts

The public ScreenSpot Pro snapshot ranks GPT-6 Astra first at 92.7%, ahead of Claude Opus 4.8 (87.9%) and GPT-5.4 (85.4%) among 18 models. We mirror the table as display-only evidence; it does not affect overall rankings.

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Use ScreenSpot Pro to judge published GUI-grounding systems after checking how each system handles the image. Direct coordinate prediction, cropping, visual search, a planner, or Python tools can change the result without changing the underlying base model.

Operator receipt: 18 sourced rows are currently displayable on this page; the leading published row is GPT-6 Astra at 92.7%.

Honest limit: ScreenSpot Pro stops at localization on a static screenshot. It does not test clicking, typing, state tracking, recovery, or completion of a multi-step workflow, and published rows with different tool or search setups are not a clean model-only comparison.

Benchmark score on ScreenSpot Pro — September 10, 2026

We mirror the published score view for ScreenSpot Pro. GPT-6 Astra leads the public snapshot at 92.7%, followed by Claude Opus 4.8 (87.9%) and GPT-5.4 (85.4%). We do not use these results to rank models overall.

18 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 10, 2026

Benchmark score table (18 models)

Score
1
GPT-6 AstraOpenAI · Closed
92.7%
2
Claude Opus 4.8Anthropic · Closed
87.9%
3
GPT-5.4OpenAI · Closed
85.4%
4
Qwen3.8 MaxAlibaba · Open weight
84.5%
5
Gemini 3.1 ProGoogle · Closed
84.4%
6
Muse SparkMeta · Closed
84.1%
7
Claude Opus 4.6Anthropic · Closed
83.1%
8
Qwen3.7 PlusAlibaba · Closed
79.0%
9
Muse Glimmer 30BMeta · Open weight
75.4%
10
Gemini 3 ProGoogle · Closed
72.7%
11
Holo2-235B-A22BH Company · Open weight
70.6%
12
Qwen3.6 PlusAlibaba · Closed
68.2%
13
Holo2-30B-A3BH Company · Open weight
66.1%
14
Qwen3.5 397BAlibaba · Open weight
65.6%
15
Holo2-8BH Company · Open weight
58.9%
16
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
57.8%
17
Holo2-4BH Company · Open weight
57.2%
18
Claude Opus 4.5Anthropic · Closed
45.7%

The published ScreenSpot Pro snapshot places GPT-6 Astra first at 92.7%. The third row is 7.3 points behind. The broader top-10 range is 20.0 points, so the table still separates the published systems.

18 models have been evaluated on ScreenSpot Pro. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. ScreenSpot Pro is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ScreenSpot Pro

Year

2025

Tasks

1,581 grounding instructions

Format

Static interface element localization

Difficulty

Professional GUI grounding

The benchmark spans 23 applications, six application categories, and three operating systems. It scores whether a system localizes the requested text or icon in a static professional screenshot, using micro-average accuracy on the official leaderboard.

BenchLM freshness & provenance

Version

ScreenSpot Pro 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ScreenSpot Pro measure?

ScreenSpot Pro measures whether a model can map a natural-language instruction to the correct target in a full-screen, high-resolution professional interface. Its 1,581 examples span 23 applications, six application categories, and three operating systems. The score is grounding accuracy, not end-to-end task completion.

What does a high ScreenSpot Pro score mean?

A high score means the published system localized more requested interface elements under its reported setup. It does not prove the same base model will click correctly inside an agent. Cropping, visual search, planners, Python tools, image resolution, and decoding policy can materially change the result.

Can ScreenSpot Pro pick the best computer-use agent?

No. ScreenSpot Pro tests static-screen grounding, which is one prerequisite for computer use. It does not test state tracking, typing, recovery after a bad action, or completion of a multi-step workflow. Pair it with OSWorld Verified and an application-specific operator test before choosing an agent.

Last updated: September 10, 2026 · BenchLM version ScreenSpot Pro 2025

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.