# OSWorld 2.0

> A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

Canonical page: https://benchlm.ai/benchmarks/osworld2

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About OSWorld 2.0

- Year: 2026
- Tasks: 108 long-horizon computer-use workflows
- Format: Interactive computer-use evaluation
- Difficulty: Long-horizon professional workflows
- Paper: [OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks](https://arxiv.org/abs/2606.29537)

OSWorld 2.0 expands computer-use evaluation to 108 long-horizon workflows that require state tracking, cross-source reasoning, visual-spatial precision, dynamic interaction, and verification. BenchLM stores the primary binary-completion score as a display-only agentic benchmark.

OSWorld 2.0 is currently weighted in BenchLM's scoring formula. The Agentic category carries 22% of the overall score, and OSWorld 2.0 contributes 10% of that category score.

## Leaderboard (20 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 72.6% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 70.6% |
| 3 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 66.9% |
| 4 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 62.6% |
| 5 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 59.0% |
| 6 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 50.2% |
| 7 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 47.9% |
| 8 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 45.6% |
| 9 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 41.7% |
| 10 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 20.6% |
| 11 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 19.4% |
| 12 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 19.4% |
| 13 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 18.2% |
| 14 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 14.2% |
| 15 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 13.9% |
| 16 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 13.0% |
| 17 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 8.3% |
| 18 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 4.6% |
| 19 | [MiniMax M3](/models/minimax-m3) | MiniMax | 4.6% |
| 20 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 2.8% |

## FAQ

### What does OSWorld 2.0 measure?

A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

### Which model scores highest on OSWorld 2.0?

GPT-6 Astra by OpenAI currently leads with a score of 72.6% on OSWorld 2.0.

### How many models are evaluated on OSWorld 2.0?

20 AI models have been evaluated on OSWorld 2.0 on BenchLM.

### Does OSWorld 2.0 affect BenchLM's overall score?

Yes. OSWorld 2.0 is a weighted benchmark inside the Agentic category, which carries 22% of BenchLM's overall score. OSWorld 2.0 itself contributes 10% of that category score.

## Compare Top Models on OSWorld 2.0

- [GPT-6 Astra vs Claude Opus 5](/compare/claude-opus-5-vs-gpt-6-astra)
- [Claude Opus 5 vs Muse Spark 1.3](/compare/claude-opus-5-vs-muse-spark-1-3)
- [Muse Spark 1.3 vs GPT-5.6 Sol](/compare/gpt-5-6-sol-vs-muse-spark-1-3)
- [GPT-5.6 Sol vs Gemini 3.8 Flash](/compare/gemini-3-8-flash-vs-gpt-5-6-sol)
