DiG-bench: Discovery in Games (DiG-bench)
Tests scientific discovery by asking agents to infer undisclosed rules across 70 interactive text-based games.
How we show DiG-bench
We mirror DiG-bench's combined scientific-discovery leaderboard. Agents play 70 interactive text games whose rules are withheld, using the same states, actions, and step budgets as human testers.
The combined score weights game wins at 80% and partial progress at 20%. Model and harness stay attached because Claude Code and the basic harness produce separate rows. Confidence bounds come from the benchmark's published table.
Expected composite score on DiG-bench — August 14, 2026
BenchLM mirrors the published expected composite score view for DiG-bench. Claude Opus 5 leads the public snapshot at 68.5% , followed by Claude Fable 5 (57.3%) and Claude Opus 4.8 (42.4%). We do not use these results to rank models overall.
Claude Opus 5
Anthropic
Basic Harness
claude-opus-5
Claude Fable 5
Anthropic
Claude Code
claude-fable-5
Claude Opus 4.8
Anthropic
Basic Harness
claude-opus-4-8
Expected composite score table (9 models)
ScoreThe published DiG-bench snapshot places Claude Opus 5 first at 68.5%. The third row is 26.1 points behind. The broader top-10 range is 62.8 points, so the table still separates the published systems.
9 models have been evaluated on DiG-bench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. DiG-bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About DiG-bench
Year
2026
Tasks
70 interactive games across seven difficulty tiers
Format
Expected composite score with confidence interval
Difficulty
Unknown-rule discovery and interactive reasoning
Humans and models receive the same game states, legal actions, and step budgets. The combined leaderboard weights successful game completion at 80% and partial progress at 20%, with model-harness variants kept separate. We mirror the score and published confidence bounds as display-only evidence.
BenchLM freshness & provenance
Version
DiG-bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does DiG-bench measure?
Tests scientific discovery by asking agents to infer undisclosed rules across 70 interactive text-based games.
Which model leads the published DiG-bench snapshot?
Claude Opus 5 currently leads the published DiG-bench snapshot with 68.5% expected composite score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on DiG-bench?
9 AI models are included in BenchLM's mirrored DiG-bench snapshot, based on the public leaderboard captured on August 14, 2026.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.