# OSWorld

> A computer-use benchmark for GUI task completion across the broader OSWorld task suite.

Canonical page: https://benchlm.ai/benchmarks/osworld

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About OSWorld

- Year: 2026
- Tasks: Computer-use tasks
- Format: Interactive GUI evaluation
- Difficulty: Broad computer-use suite
- Paper: [GLM-5V-Turbo](https://docs.z.ai/guides/vlm/glm-5v-turbo)

BenchLM tracks plain OSWorld as a display-only provider-table reference and preserves OSWorld-Verified as the weighted core benchmark key.

OSWorld is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 66.3% |
| 2 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 47.4% |

## FAQ

### What does OSWorld measure?

A computer-use benchmark for GUI task completion across the broader OSWorld task suite.

### Which model scores highest on OSWorld?

Claude Opus 4.5 by Anthropic currently leads with a score of 66.3% on OSWorld.

### How many models are evaluated on OSWorld?

2 AI models have been evaluated on OSWorld on BenchLM.

### Does OSWorld affect BenchLM's overall score?

Not directly. OSWorld is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on OSWorld

- [Claude Opus 4.5 vs Nemotron 3 Nano Omni 30B A3B](/compare/claude-opus-4-5-vs-nemotron-3-nano-omni-30b-a3b)
