Skip to main content

Benchmark profile

OSWorld 2.0

A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

Data verified

Benchmark score on OSWorld 2.0 — July 29, 2026

BenchLM mirrors the published score view for OSWorld 2.0. Claude Opus 5 leads the public snapshot at 70.6% , followed by GPT-5.6 Sol (62.6%) and GPT-5.6 Terra (50.2%). BenchLM does not use these results to rank models overall.

13 modelsAgenticCurrentDisplay onlyUpdated July 29, 2026

Benchmark score table (13 models)

Score
1
Claude Opus 5Anthropic · Closed
70.6%
2
GPT-5.6 SolOpenAI · Closed
62.6%
3
GPT-5.6 TerraOpenAI · Closed
50.2%
4
GPT-5.6 LunaOpenAI · Closed
45.6%
5
Claude Opus 4.8Anthropic · Closed
20.6%
6
Claude Opus 4.7 (Adaptive)Anthropic · Closed
18.2%
7
Muse Spark 1.1Meta · Closed
14.2%
8
Claude Opus 4.7Anthropic · Closed
13.9%
9
GPT-5.5OpenAI · Closed
13.0%
10
Claude Sonnet 4.6Anthropic · Closed
8.3%
11
Kimi K2.6Moonshot AI · Open weight
4.6%
12
MiniMax M3MiniMax · Open weight
4.6%
13
Qwen3.7 PlusAlibaba · Closed
2.8%

The published OSWorld 2.0 snapshot places Claude Opus 5 first at 70.6%. The third row is 20.4 points behind. The broader top-10 range is 62.3 points, so the table still separates the published systems.

13 models have been evaluated on OSWorld 2.0. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. OSWorld 2.0 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About OSWorld 2.0

Year

2026

Tasks

108 long-horizon computer-use workflows

Format

Interactive computer-use evaluation

Difficulty

Long-horizon professional workflows

OSWorld 2.0 expands computer-use evaluation to 108 long-horizon workflows that require state tracking, cross-source reasoning, visual-spatial precision, dynamic interaction, and verification. BenchLM stores the primary binary-completion score as a display-only agentic benchmark.

BenchLM freshness & provenance

Version

OSWorld 2.0 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does OSWorld 2.0 measure?

A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

Which model scores highest on OSWorld 2.0?

Claude Opus 5 by Anthropic currently leads with a score of 70.6% on OSWorld 2.0.

How many models are evaluated on OSWorld 2.0?

13 AI models have been evaluated on OSWorld 2.0 on BenchLM.

Last updated: July 29, 2026 · BenchLM version OSWorld 2.0 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.