# Vals BioMysteryBench agent leaderboard (BioMysteryBench (Vals))

> Independent agent runs on biological data investigations, split by human-solvable and human-difficult tasks.

Canonical page: https://benchlm.ai/benchmarks/vals-biomysterybench

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About BioMysteryBench (Vals)

- Year: 2026
- Tasks: End-to-end bioinformatics investigations from raw data
- Format: Average task accuracy across three runs
- Difficulty: Agentic biological analysis
- Paper: [Vals BioMysteryBench](https://www.vals.ai/benchmarks/biomysterybench)

Vals runs Anthropic's biological-data tasks through a fixed Terminus 2 harness and reports overall, human-solvable, and human-difficult scores. Three full-benchmark runs and their uncertainty remain source-specific, display-only evidence.

BioMysteryBench (Vals) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (19 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Sonnet 5.5](/models/claude-sonnet-5-5) | Terminus 2 · Terminus 2 | Anthropic | 81.11% |
| 2 | [GPT-6 Astra](/models/gpt-6-astra) | Terminus 2 · max reasoning · Terminus 2 | OpenAI | 79.26% |
| 3 | [Claude Opus 5.5](/models/claude-opus-5-5) | Terminus 2 · Terminus 2 | Anthropic | 79.26% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | Terminus 2 · Terminus 2 | Anthropic | 79.26% |
| 5 | [GPT-6 Sol](/models/gpt-6-sol) | Terminus 2 · max reasoning · Terminus 2 | OpenAI | 74.81% |
| 6 | [Grok 4.6](/models/grok-4-6) | Terminus 2 · high reasoning · Terminus 2 | xAI | 72.22% |
| 7 | [Kimi K3](/models/kimi-k3) | Terminus 2 · max reasoning · Terminus 2 | Moonshot AI | 71.48% |
| 8 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | Terminus 2 · max reasoning · Terminus 2 | OpenAI | 71.11% |
| 9 | [Grok 4.7](/models/grok-4-7) | Terminus 2 · xhigh reasoning · Terminus 2 | xAI | 69.26% |
| 10 | [Hy4 preview](/models/hy4-preview) | Terminus 2 · Terminus 2 | Tencent | 69.26% |
| 11 | [MiMo-V2.6-Flash](/models/mimo-v2-6-flash) | Terminus 2 · Terminus 2 | Xiaomi | 69.26% |
| 12 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | Terminus 2 · high reasoning · Terminus 2 | DeepSeek | 67.78% |
| 13 | [Muse Spark 1.2](/models/muse-spark-1-2) | Terminus 2 · xhigh reasoning · Terminus 2 | Meta | 64.81% |
| 14 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | Terminus 2 · high reasoning · Terminus 2 | DeepSeek | 64.44% |
| 15 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Terminus 2 · high reasoning · Terminus 2 | Google | 62.22% |
| 16 | [GPT-6 Luna](/models/gpt-6-luna) | Terminus 2 · max reasoning · Terminus 2 | OpenAI | 61.48% |
| 17 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | Terminus 2 · max reasoning · Terminus 2 | OpenAI | 61.48% |
| 18 | [Gemini 3.1 Pro Preview](https://www.vals.ai/models/google_gemini-3.1-pro-preview) | Terminus 2 · high reasoning · Terminus 2 | Google | 60.00% |
| 19 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | Terminus 2 · high reasoning · Terminus 2 | Google | 58.52% |

## FAQ

### What does BioMysteryBench (Vals) measure?

Independent agent runs on biological data investigations, split by human-solvable and human-difficult tasks.

### Which model leads the published BioMysteryBench (Vals) snapshot?

Claude Sonnet 5.5 currently leads the published BioMysteryBench (Vals) snapshot with a score of 81.11%.

### How many models are evaluated on BioMysteryBench (Vals)?

The September 27, 2026 contains 19 AI models.

### Does BioMysteryBench (Vals) affect BenchLM's overall score?

Not directly. BioMysteryBench (Vals) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
