# GeneBench-Pro

> A multistage statistical-reasoning benchmark for genomics and biological-data analysis agents.

Canonical page: https://benchlm.ai/benchmarks/genebenchpro

- Category: [Reasoning](/reasoning)
- Last updated: September 10, 2026

## About GeneBench-Pro

- Year: 2026
- Tasks: 129 genomics statistical-analysis workflows
- Format: Eval-level pass rate across dependent analysis decisions
- Difficulty: Long-horizon scientific reasoning
- Paper: [GeneBench-Pro: Evaluating Multistage Statistical Reasoning](https://cdn.openai.com/pdf/21938268-21af-442f-af93-3b2249afb241/genebench-pro.pdf)

GeneBench-Pro evaluates whether an agent chooses and executes the correct sequence of statistical-analysis decisions in genomics workflows. BenchLM stores the published result as display-only because it is a newly introduced, provider-developed benchmark with sparse public coverage.

GeneBench-Pro is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 37.8% |
| 2 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 28.7% |

## FAQ

### What does GeneBench-Pro measure?

A multistage statistical-reasoning benchmark for genomics and biological-data analysis agents.

### Which model scores highest on GeneBench-Pro?

GPT-6 Astra by OpenAI currently leads with a score of 37.8% on GeneBench-Pro.

### How many models are evaluated on GeneBench-Pro?

2 AI models have been evaluated on GeneBench-Pro on BenchLM.

### Does GeneBench-Pro affect BenchLM's overall score?

Not directly. GeneBench-Pro is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on GeneBench-Pro

- [GPT-6 Astra vs GPT-5.6 Sol](/compare/gpt-5-6-sol-vs-gpt-6-astra)
