Skip to main content
BenchLM

DiG-bench: Discovery in Games (DiG-bench)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Tests scientific discovery by asking agents to infer undisclosed rules across 70 interactive text-based games.

Expected composite score on DiG-bench — September 29, 2026

We mirror the published expected composite score view for DiG-bench. Claude Opus 5 leads the public snapshot at 68.5%, followed by Claude Fable 5 (57.3%) and Claude Opus 4.8 (42.4%). We do not use these results to rank models overall.

9 modelsAgenticCurrentDisplay onlyUpdated September 29, 2026

Expected composite score table (9 models)

Score
1
Claude Opus 5Anthropic · ClosedBasic HarnessBasic Harness
68.5%
2
Claude Fable 5Anthropic · ClosedClaude CodeClaude Code
57.3%
3
Claude Opus 4.8Anthropic · ClosedBasic HarnessBasic Harness
42.4%
4
Claude Opus 4.8Anthropic · ClosedClaude CodeClaude Code
31.4%
5
GPT-5.5OpenAI · ClosedBasic HarnessBasic Harness
27.9%
6
Gemini 3.1 ProGoogle · ClosedBasic HarnessBasic Harness
19.4%
7
GLM-5.2Z.AI · Open weightBasic HarnessBasic Harness
18.6%
8
Qwen3.5 397BAlibaba · Open weightBasic HarnessBasic Harness
5.9%
9
DeepSeek V4 Pro 0813DeepSeek · Open weightBasic HarnessBasic Harness
5.7%

How we show DiG-bench

We mirror DiG-bench's combined scientific-discovery leaderboard. Agents play 70 interactive text games whose rules are withheld, using the same states, actions, and step budgets as human testers.

The combined score weights game wins at 80% and partial progress at 20%. Model and harness stay attached because Claude Code and the basic harness produce separate rows. Confidence bounds come from the benchmark's published table.

The published DiG-bench snapshot places Claude Opus 5 first at 68.5%. The third row is 26.1 points behind. The broader top-10 range is 62.8 points, so the table still separates the published systems.

9 models have been evaluated on DiG-bench. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. DiG-bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About DiG-bench

Year

2026

Tasks

70 interactive games across seven difficulty tiers

Format

Expected composite score with confidence interval

Difficulty

Unknown-rule discovery and interactive reasoning

Humans and models receive the same game states, legal actions, and step budgets. The combined leaderboard weights successful game completion at 80% and partial progress at 20%, with model-harness variants kept separate. We mirror the score and published confidence bounds as display-only evidence.

Freshness and provenance

Version

DiG-bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does DiG-bench measure?

Tests scientific discovery by asking agents to infer undisclosed rules across 70 interactive text-based games.

Which model leads the published DiG-bench snapshot?

Claude Opus 5 currently leads the published DiG-bench snapshot with 68.5% expected composite score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on DiG-bench?

The September 29, 2026 snapshot contains 9 AI models.

Last updated: September 29, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.