Skip to main content
BenchLM

Vals CUA-bench (CUA-bench)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A Vals AI computer-use benchmark in which agents play six commercial video games with a keyboard and mouse.

CUA-bench score on CUA-bench — September 22, 2026

We mirror the published cua-bench score view for CUA-bench. GPT-6 Astra leads the public snapshot at 19.17%, followed by Claude Opus 5.5 (14.00%) and Claude Fable 5.1 (13.17%). We do not use these results to rank models overall.

7 modelsAgenticCurrentDisplay onlyUpdated September 22, 2026

CUA-bench score table (7 models)

Score
1
GPT-6 AstraOpenAI · Closedcodex · max reasoningcodex
19.17%
2
Claude Opus 5.5Anthropic · Closedclaude-code · max reasoningclaude-code
14.00%
3
Claude Fable 5.1Anthropic · Closedclaude-code · max reasoningclaude-code
13.17%
4
Claude Opus 5Anthropic · Closedclaude-code · max reasoningclaude-code
9.00%
5
GPT-5.6 SolOpenAI · Closedcodex · max reasoningcodex
8.33%
6
Muse Spark 1.3 Maxmetamuse-code · max reasoningmuse-code
5.83%
7
Gemini 3.8 FlashGoogle · Closedgoogle-computer-use · high reasoninggoogle-computer-use
4.17%

How CUA-bench is shown here

BenchLM mirrors the public Vals AI CUA-bench leaderboard captured from https://www.vals.ai/benchmarks/cua_bench and updated by Vals on September 22, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

CUA-bench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

7 Vals rows7 task viewsprivate datasetTasks: Overall, Minecraft, Hidden sandbox, SUPERHOT, Hidden reflexDisplay only

The published CUA-bench snapshot places GPT-6 Astra first at 19.17%. The third row is 6.00 points behind. The broader top-10 range is 15.00 points, so the table still separates the published systems.

7 models have been evaluated on CUA-bench. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. CUA-bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About CUA-bench

Year

2026

Tasks

Six commercial video-game control tasks, including held-out games

Format

Overall score with game-level task splits

Difficulty

Keyboard-and-mouse computer use in real-time games

The public beta table reports overall and task-level scores for Minecraft, SUPERHOT, eFootball, and three held-out games. Each result pairs a model with a provider-specific computer-use harness, while the task set remains private, so this board is display-only and excluded from weighted model rankings.

Freshness and provenance

Version

CUA-bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CUA-bench measure?

A Vals AI computer-use benchmark in which agents play six commercial video games with a keyboard and mouse.

Which model leads the published CUA-bench snapshot?

GPT-6 Astra currently leads the published CUA-bench snapshot with 19.17% cua-bench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on CUA-bench?

The September 22, 2026 snapshot contains 7 AI models.

Last updated: September 22, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.