GameDevBench
Evaluates coding agents on 333 multimodal game-development tasks in Godot, spanning 2D graphics, 3D graphics, user interfaces, and gameplay logic.
How to read this leaderboard
Higher Pass@1 means the published model-and-agent setup completed more of the 333 tasks. Read the harness, reasoning-effort label, and confidence interval with the score: several model families appear more than once because the authors tested distinct configurations.
Operator receipt: 17 sourced rows are currently displayable on this page; the leading published row is Claude Fable 5 at 67.3%.
Honest limit: GameDevBench does not isolate the base model. Agent harnesses, visual feedback, reasoning effort, and environment setup all change the result, while overlapping 95% confidence intervals make small score gaps uncertain. We therefore keep the leaderboard display only.
How we show GameDevBench
We mirror the official GameDevBench ICML 2026 camera-ready leaderboard across 333 Godot tasks drawn from 88 tutorials. The published table reports Pass@1 for 17 model-and-agent configurations and includes a 95% confidence interval for every row.
GameDevBench covers 4 skill groups: 2D graphics, 3D graphics, user interface work, and gameplay logic. Each public score uses the model's best published harness and multimodal-feedback configuration, so the table preserves harness and reasoning-effort labels instead of collapsing them into a base-model result.
We keep the raw results display only. Visual feedback, agent scaffolding, reasoning effort, and the Godot environment all affect task completion, which makes the benchmark useful for comparing complete game-development setups but unsuitable for a weighted model-only rank.
Snapshot
Pass@1 on GameDevBench ICML 2026 camera-ready leaderboard — ICML 2026 camera-ready results, updated August 13, 2026
BenchLM mirrors the published pass@1 view for GameDevBench ICML 2026 camera-ready leaderboard. Claude Fable 5 leads the public snapshot at 67.3% , followed by GPT-5.6 Sol (63.7%) and GPT-5.6 Sol (63.1%). We do not use these results to rank models overall.
Claude Fable 5
Anthropic
Claude Code
Claude Code · xhigh · 95% CI ±5
GPT-5.6 Sol
OpenAI
Codex
Codex · xhigh · 95% CI ±5.2
GPT-5.6 Sol
OpenAI
Codex
Codex · high · 95% CI ±5.2
Pass@1 table (17 models)
ScoreThe published GameDevBench snapshot places Claude Fable 5 first at 67.3%. The third row is 4.2 points behind. The broader top-10 range is 15.3 points, so the table still separates the published systems.
17 models have been evaluated on GameDevBench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. GameDevBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About GameDevBench
Year
2026
Tasks
333 tasks from 88 tutorials
Format
Pass@1 on the full task set with 95% confidence intervals
Difficulty
Multimodal game development in Godot 4.4.1
GameDevBench derives 333 tasks from 88 web and video tutorials and runs them in Godot 4.4.1. The camera-ready leaderboard reports Pass@1 on the full task set using each model's best published agent harness and multimodal-feedback configuration, with 95% confidence intervals kept alongside the scores.
BenchLM freshness & provenance
Version
GameDevBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does GameDevBench measure?
GameDevBench measures whether coding agents can complete 333 game-development tasks in Godot. Tasks cover 2D graphics, 3D graphics, user interfaces, and gameplay logic, requiring agents to work with code and visual assets such as shaders, sprites, animations, and scenes.
Which configuration leads GameDevBench?
Claude Fable 5 at xhigh effort in Claude Code leads the August 13, 2026 camera-ready table at 67.3% Pass@1, with a reported 95% confidence interval of ±5.0 points. GPT-5.6 Sol at xhigh effort in Codex follows at 63.7% ±5.2.
Does GameDevBench affect BenchLM rankings?
No. We keep GameDevBench display only because each row combines a model with an agent harness, reasoning setting, multimodal-feedback setup, and Godot environment. Those choices are part of the measured system, so the scores are not normalized base-model results.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.