# ScreenSpot Pro

> A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.

Canonical page: https://benchlm.ai/benchmarks/screenspot-pro

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 10, 2026

## About ScreenSpot Pro

- Year: 2025
- Tasks: 1,581 grounding instructions
- Format: Static interface element localization
- Difficulty: Professional GUI grounding
- Paper: [ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use](https://arxiv.org/abs/2504.07981)

The benchmark spans 23 applications, six application categories, and three operating systems. It scores whether a system localizes the requested text or icon in a static professional screenshot, using micro-average accuracy on the official leaderboard.

ScreenSpot Pro is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (18 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 92.7% |
| 2 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 87.9% |
| 3 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 85.4% |
| 4 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 84.5% |
| 5 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 84.4% |
| 6 | [Muse Spark](/models/muse-spark) | Meta | 84.1% |
| 7 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 83.1% |
| 8 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 79.0% |
| 9 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 75.4% |
| 10 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 72.7% |
| 11 | [Holo2-235B-A22B](/models/holo2-235b-a22b) | H Company | 70.6% |
| 12 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 68.2% |
| 13 | [Holo2-30B-A3B](/models/holo2-30b-a3b) | H Company | 66.1% |
| 14 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 65.6% |
| 15 | [Holo2-8B](/models/holo2-8b) | H Company | 58.9% |
| 16 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 57.8% |
| 17 | [Holo2-4B](/models/holo2-4b) | H Company | 57.2% |
| 18 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 45.7% |

## FAQ

### What does ScreenSpot Pro measure?

ScreenSpot Pro measures whether a model can map a natural-language instruction to the correct target in a full-screen, high-resolution professional interface. Its 1,581 examples span 23 applications, six application categories, and three operating systems. The score is grounding accuracy, not end-to-end task completion.

### What does a high ScreenSpot Pro score mean?

A high score means the published system localized more requested interface elements under its reported setup. It does not prove the same base model will click correctly inside an agent. Cropping, visual search, planners, Python tools, image resolution, and decoding policy can materially change the result.

### Can ScreenSpot Pro pick the best computer-use agent?

No. ScreenSpot Pro tests static-screen grounding, which is one prerequisite for computer use. It does not test state tracking, typing, recovery after a bad action, or completion of a multi-step workflow. Pair it with OSWorld Verified and an application-specific operator test before choosing an agent.

## Compare Top Models on ScreenSpot Pro

- [GPT-6 Astra vs Claude Opus 4.8](/compare/claude-opus-4-8-vs-gpt-6-astra)
- [Claude Opus 4.8 vs GPT-5.4](/compare/claude-opus-4-8-vs-gpt-5-4)
- [GPT-5.4 vs Qwen3.8 Max](/compare/gpt-5-4-vs-qwen3-8-max)
- [Qwen3.8 Max vs Gemini 3.1 Pro](/compare/gemini-3-1-pro-vs-qwen3-8-max)
