# DiG-bench: Discovery in Games (DiG-bench)

> Tests scientific discovery by asking agents to infer undisclosed rules across 70 interactive text-based games.

Canonical page: https://benchlm.ai/benchmarks/digbench

- Category: [Agentic](/agentic)
- Last updated: September 29, 2026

## About DiG-bench

- Year: 2026
- Tasks: 70 interactive games across seven difficulty tiers
- Format: Expected composite score with confidence interval
- Difficulty: Unknown-rule discovery and interactive reasoning
- Paper: [DiG-bench: Discovery in Games](https://digbench.ai/)

Humans and models receive the same game states, legal actions, and step budgets. The combined leaderboard weights successful game completion at 80% and partial progress at 20%, with model-harness variants kept separate. We mirror the score and published confidence bounds as display-only evidence.

DiG-bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (9 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Basic Harness · Basic Harness | Anthropic | 68.5% |
| 2 | [Claude Fable 5](/models/claude-fable) | Claude Code · Claude Code | Anthropic | 57.3% |
| 3 | [Claude Opus 4.8](/models/claude-opus-4-8) | Basic Harness · Basic Harness | Anthropic | 42.4% |
| 4 | [Claude Opus 4.8](/models/claude-opus-4-8) | Claude Code · Claude Code | Anthropic | 31.4% |
| 5 | [GPT-5.5](/models/gpt-5-5) | Basic Harness · Basic Harness | OpenAI | 27.9% |
| 6 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Basic Harness · Basic Harness | Google | 19.4% |
| 7 | [GLM-5.2](/models/glm-5-2) | Basic Harness · Basic Harness | Z.AI | 18.6% |
| 8 | [Qwen3.5 397B](/models/qwen3-5-397b) | Basic Harness · Basic Harness | Alibaba | 5.9% |
| 9 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | Basic Harness · Basic Harness | DeepSeek | 5.7% |

## FAQ

### What does DiG-bench measure?

Tests scientific discovery by asking agents to infer undisclosed rules across 70 interactive text-based games.

### Which model leads the published DiG-bench snapshot?

Claude Opus 5 currently leads the published DiG-bench snapshot with a score of 68.5%.

### How many models are evaluated on DiG-bench?

The September 29, 2026 contains 9 AI models.

### Does DiG-bench affect BenchLM's overall score?

Not directly. DiG-bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
