Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

DiG-bench: Discovery in Games (DiG-bench)

Tests scientific discovery by asking agents to infer undisclosed rules across 70 interactive text-based games.

Data verified 28 confirmed releases in the last 30 daysStart free brief

How we show DiG-bench

We mirror DiG-bench's combined scientific-discovery leaderboard. Agents play 70 interactive text games whose rules are withheld, using the same states, actions, and step budgets as human testers.

The combined score weights game wins at 80% and partial progress at 20%. Model and harness stay attached because Claude Code and the basic harness produce separate rows. Confidence bounds come from the benchmark's published table.

9 model-harness rows70 games767 recorded runsDisplay only

Expected composite score on DiG-bench — August 14, 2026

BenchLM mirrors the published expected composite score view for DiG-bench. Claude Opus 5 leads the public snapshot at 68.5% , followed by Claude Fable 5 (57.3%) and Claude Opus 4.8 (42.4%). We do not use these results to rank models overall.

9 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated August 14, 2026

Expected composite score table (9 models)

Score
1
Claude Opus 5Anthropic · ClosedBasic Harness
68.5%
2
Claude Fable 5Anthropic · ClosedClaude Code
57.3%
3
Claude Opus 4.8Anthropic · ClosedBasic Harness
42.4%
4
Claude Opus 4.8Anthropic · ClosedClaude Code
31.4%
5
GPT-5.5OpenAI · ClosedBasic Harness
27.9%
6
Gemini 3.1 ProGoogle · ClosedBasic Harness
19.4%
7
GLM-5.2Z.AI · Open weightBasic Harness
18.6%
8
Qwen3.5 397BAlibaba · Open weightBasic Harness
5.9%
9
DeepSeek V4 ProDeepSeekBasic Harness
5.7%

The published DiG-bench snapshot places Claude Opus 5 first at 68.5%. The third row is 26.1 points behind. The broader top-10 range is 62.8 points, so the table still separates the published systems.

9 models have been evaluated on DiG-bench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. DiG-bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About DiG-bench

Year

2026

Tasks

70 interactive games across seven difficulty tiers

Format

Expected composite score with confidence interval

Difficulty

Unknown-rule discovery and interactive reasoning

Humans and models receive the same game states, legal actions, and step budgets. The combined leaderboard weights successful game completion at 80% and partial progress at 20%, with model-harness variants kept separate. We mirror the score and published confidence bounds as display-only evidence.

BenchLM freshness & provenance

Version

DiG-bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does DiG-bench measure?

Tests scientific discovery by asking agents to infer undisclosed rules across 70 interactive text-based games.

Which model leads the published DiG-bench snapshot?

Claude Opus 5 currently leads the published DiG-bench snapshot with 68.5% expected composite score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on DiG-bench?

9 AI models are included in BenchLM's mirrored DiG-bench snapshot, based on the public leaderboard captured on August 14, 2026.

Last updated: August 14, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.