# Vals CUA-bench (CUA-bench)

> A Vals AI computer-use benchmark in which agents play six commercial video games with a keyboard and mouse.

Canonical page: https://benchlm.ai/benchmarks/vals-cua-bench

- Category: [Agentic](/agentic)
- Last updated: September 22, 2026

## About CUA-bench

- Year: 2026
- Tasks: Six commercial video-game control tasks, including held-out games
- Format: Overall score with game-level task splits
- Difficulty: Keyboard-and-mouse computer use in real-time games
- Paper: [CUA-bench](https://www.vals.ai/benchmarks/cua_bench)

The public beta table reports overall and task-level scores for Minecraft, SUPERHOT, eFootball, and three held-out games. Each result pairs a model with a provider-specific computer-use harness, while the task set remains private, so this board is display-only and excluded from weighted model rankings.

CUA-bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (7 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | codex · max reasoning · codex | OpenAI | 19.17% |
| 2 | [Claude Opus 5.5](/models/claude-opus-5-5) | claude-code · max reasoning · claude-code | Anthropic | 14.00% |
| 3 | [Claude Fable 5.1](/models/claude-fable-5-1) | claude-code · max reasoning · claude-code | Anthropic | 13.17% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | claude-code · max reasoning · claude-code | Anthropic | 9.00% |
| 5 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | codex · max reasoning · codex | OpenAI | 8.33% |
| 6 | [Muse Spark 1.3 Max](https://www.vals.ai/models/meta_muse_spark_1_3_max) | muse-code · max reasoning · muse-code | meta | 5.83% |
| 7 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | google-computer-use · high reasoning · google-computer-use | Google | 4.17% |

## FAQ

### What does CUA-bench measure?

A Vals AI computer-use benchmark in which agents play six commercial video games with a keyboard and mouse.

### Which model leads the published CUA-bench snapshot?

GPT-6 Astra currently leads the published CUA-bench snapshot with a score of 19.17%.

### How many models are evaluated on CUA-bench?

The September 22, 2026 contains 7 AI models.

### Does CUA-bench affect BenchLM's overall score?

Not directly. CUA-bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
