GameDevBench
We show this table for reference; we do not rank on it.
Evaluates coding agents on 333 multimodal game-development tasks in Godot, spanning 2D graphics, 3D graphics, user interfaces, and gameplay logic.
Pass@1 on GameDevBench ICML 2026 camera-ready leaderboard — ICML 2026 camera-ready results, updated September 09, 2026
We mirror the published pass@1 view for GameDevBench ICML 2026 camera-ready leaderboard. GPT-6 Astra leads the public snapshot at 68.8%, followed by Claude Fable 5 (67.3%) and GPT-5.6 Sol (63.7%). We do not use these results to rank models overall.
GPT-6 Astra
OpenAI
Codex
Codex · high · 95% CI ±5
Claude Fable 5
Anthropic
Claude Code
Claude Code · xhigh · 95% CI ±5
GPT-5.6 Sol
OpenAI
Codex
Codex · xhigh · 95% CI ±5.2
18 modelsCodingCurrentDisplay onlyUpdated ICML 2026 camera-ready results, updated September 09, 2026
Pass@1 table (18 models)
ScoreHow to read this leaderboard
Higher Pass@1 means the published model-and-agent setup completed more of the 333 tasks. Read the harness, reasoning-effort label, and confidence interval with the score: several model families appear more than once because the authors tested distinct configurations.
Operator receipt: 18 sourced rows are currently displayable on this page; the leading published row is GPT-6 Astra at 68.8%.
Honest limit: GameDevBench does not isolate the base model. Agent harnesses, visual feedback, reasoning effort, and environment setup all change the result, while overlapping 95% confidence intervals make small score gaps uncertain. We therefore keep the leaderboard display only.
How we show GameDevBench
We mirror the official GameDevBench ICML 2026 camera-ready leaderboard across 333 Godot tasks drawn from 88 tutorials. The published table reports Pass@1 for 18 model-and-agent configurations and includes a 95% confidence interval for every row.
GameDevBench covers 4 skill groups: 2D graphics, 3D graphics, user interface work, and gameplay logic. Each public score uses the model's best published harness and multimodal-feedback configuration, so the table preserves harness and reasoning-effort labels instead of collapsing them into a base-model result.
We keep the raw results display only. Visual feedback, agent scaffolding, reasoning effort, and the Godot environment all affect task completion, which makes the benchmark useful for comparing complete game-development setups but unsuitable for a weighted model-only rank.
Snapshot
The published GameDevBench snapshot places GPT-6 Astra first at 68.8%. The third row is 5.1 points behind. The broader top-10 range is 15.0 points, so the table still separates the published systems.
18 models have been evaluated on GameDevBench. The benchmark falls in the Coding category. GameDevBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About GameDevBench
Year
2026
Tasks
333 tasks from 88 tutorials
Format
Pass@1 on the full task set with 95% confidence intervals
Difficulty
Multimodal game development in Godot 4.4.1
GameDevBench derives 333 tasks from 88 web and video tutorials and runs them in Godot 4.4.1. The camera-ready leaderboard reports Pass@1 on the full task set using each model's best published agent harness and multimodal-feedback configuration, with 95% confidence intervals kept alongside the scores.
Freshness and provenance
Version
GameDevBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does GameDevBench measure?
GameDevBench measures whether coding agents can complete 333 game-development tasks in Godot. Tasks cover 2D graphics, 3D graphics, user interfaces, and gameplay logic, requiring agents to work with code and visual assets such as shaders, sprites, animations, and scenes.
Which configuration leads GameDevBench?
Claude Fable 5 at xhigh effort in Claude Code leads the August 13, 2026 camera-ready table at 67.3% Pass@1, with a reported 95% confidence interval of ±5.0 points. GPT-5.6 Sol at xhigh effort in Codex follows at 63.7% ±5.2.
Does GameDevBench affect BenchLM rankings?
No. We keep GameDevBench display only because each row combines a model with an agent harness, reasoning setting, multimodal-feedback setup, and Godot environment. Those choices are part of the measured system, so the scores are not normalized base-model results.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.