# OSWorld-Verified

> OSWorld-Verified is the July 2025 repaired release of OSWorld's real-computer evaluation. It measures whether a model-agent system can finish desktop and web tasks from configured starting states, with success checked by execution-based evaluators.

Canonical page: https://benchlm.ai/benchmarks/osworld-verified

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About OSWorld-Verified

- Year: 2025
- Tasks: 369 real-world computer tasks (361 when eight Google Drive tasks are excluded)
- Format: Execution-based interactive task success
- Difficulty: Multi-step desktop and cross-application workflows
- Paper: [OSWorld](https://os-world.github.io/)

The release covers 369 real-world tasks across desktop and web applications. The maintainers allow eight Google Drive tasks to be manually configured or excluded, making a 361-task run officially acceptable. Public evaluation requires the maintainers to run the agent or review monitoring data and trajectories. OSWorld 2.0 is a newer, separate protocol.

OSWorld-Verified is currently weighted in BenchLM's scoring formula. The Agentic category carries 22% of the overall score, and OSWorld-Verified contributes 25% of that category score.

## Leaderboard (32 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 86.1% |
| 2 | [Claude Fable 5](/models/claude-fable) | Anthropic | 85% |
| 3 | [Claude Mythos 5](/models/claude-mythos-5) | Anthropic | 85% |
| 4 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 84.3% |
| 5 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 83.4% |
| 6 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | Google | 83% |
| 7 | [Holo3-35B-A3B](/models/holo3-35b-a3b) | H Company | 82.6% |
| 8 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 81.2% |
| 9 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 80.8% |
| 10 | [Holo3-122B-A10B](/models/holo3-122b-a10b) | H Company | 78.8% |
| 11 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 78.7% |
| 12 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 78.4% |
| 13 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 78% |
| 14 | [UI-Mate-27B](/models/ui-mate-27b) | Tencent | 77% |
| 15 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 75% |
| 16 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | Google | 74% |
| 17 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 73.3% |
| 18 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 73.1% |
| 19 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 72.7% |
| 20 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 72.1% |
| 21 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 72.1% |
| 22 | [MiniMax M3](/models/minimax-m3) | MiniMax | 70.1% |
| 23 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 66.3% |
| 24 | [UI-Mate-9B](/models/ui-mate-9b) | Tencent | 66.2% |
| 25 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 65.9% |
| 26 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 64.7% |
| 27 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 61.4% |
| 28 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 58% |
| 29 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 56.2% |
| 30 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 54.5% |
| 31 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 47.3% |
| 32 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenAI | 39% |

## FAQ

### What does OSWorld-Verified measure?

OSWorld-Verified measures whether a computer-use agent can complete 369 tasks across real desktop and web applications. Each task starts from a configured state and is checked with an execution-based evaluator. Eight Google Drive tasks may be manually configured or excluded, producing an officially accepted 361-task run.

### Are OSWorld-Verified scores directly comparable?

Only when the setup matches. Compare the same task set, environment revision, model-agent variant, observation and action interface, prompt or scaffold, action budget, and attempt policy. The official leaderboard separates general models, specialized models, and agentic frameworks; provider-published scores may use different conditions.

### Can OSWorld-Verified choose the best computer-use model?

Not alone. It is strong evidence for desktop task completion, but fixed applications cannot reproduce every production login, permission, network, app-version, latency, cost, or safety condition. Use matched OSWorld-Verified results with workflow trials, and treat OSWorld 2.0 as a separate newer protocol rather than interchangeable evidence.

## Compare Top Models on OSWorld-Verified

- [Qwen3.8 Max vs Claude Fable 5](/compare/claude-fable-vs-qwen3-8-max)
- [Claude Fable 5 vs Claude Mythos 5](/compare/claude-fable-vs-claude-mythos-5)
- [Claude Mythos 5 vs Qwen3.8-27B](/compare/claude-mythos-5-vs-qwen3-8-27b)
- [Qwen3.8-27B vs Claude Opus 4.8](/compare/claude-opus-4-8-vs-qwen3-8-27b)

## Related Reading

- [OSWorld-Verified benchmark explainer](/blog/posts/osworld-verified-computer-use-benchmark)
