# Vals MysteryMechanism (MysteryMechanism)

> Can agents rediscover sealed mathematical mechanisms through bounded experiments?

Canonical page: https://benchlm.ai/benchmarks/mysterymechanism

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About MysteryMechanism

- Year: 2026
- Tasks: Sealed mathematical mechanisms probed under an experiment budget
- Format: Accuracy score
- Difficulty: Agentic scientific discovery
- Paper: [Vals MysteryMechanism](https://www.vals.ai/benchmarks/mysterymechanism)

An agent is given a sealed mechanism it can only probe through a bounded budget of experiments, and is scored on how well it recovers the underlying rule. The task set is private and the runner is external, so the board stays display only and outside weighted rankings.

MysteryMechanism is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (19 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 53.15% |
| 2 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 49.55% |
| 3 | [Claude Sonnet 5.5](/models/claude-sonnet-5-5) | Anthropic | 49.10% |
| 4 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 47.75% |
| 5 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 37.39% |
| 6 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 36.49% |
| 7 | [Muse Spark 1.3 Max](https://www.vals.ai/models/meta_muse_spark_1_3_max) | Meta | 36.04% |
| 8 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 33.33% |
| 9 | [Grok 4.6](/models/grok-4-6) | xAI | 30.63% |
| 10 | [GPT-6 Sol](/models/gpt-6-sol) | OpenAI | 30.18% |
| 11 | [Grok 4.7](/models/grok-4-7) | xAI | 25.23% |
| 12 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 23.87% |
| 13 | [GLM-5.3](/models/glm-5-3) | Z.AI | 22.97% |
| 14 | [MiMo-V2.6-Flash](/models/mimo-v2-6-flash) | Xiaomi | 21.62% |
| 15 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 21.17% |
| 16 | [Gemini 3.1 Pro Preview](https://www.vals.ai/models/google_gemini-3.1-pro-preview) | Google | 20.72% |
| 17 | [GPT-6 Luna](/models/gpt-6-luna) | OpenAI | 19.37% |
| 18 | [MiMo-V2.6-Pro](/models/mimo-v2-6-pro) | Xiaomi | 15.31% |
| 19 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 14.41% |

## FAQ

### What does MysteryMechanism measure?

Can agents rediscover sealed mathematical mechanisms through bounded experiments?

### Which model leads the published MysteryMechanism snapshot?

GPT-6 Astra currently leads the published MysteryMechanism snapshot with a score of 53.15%.

### How many models are evaluated on MysteryMechanism?

The September 27, 2026 contains 19 AI models.

### Does MysteryMechanism affect BenchLM's overall score?

Not directly. MysteryMechanism is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
