# Scale Labs DrugDiscoveryBench (DrugDiscoveryBench)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-drugdiscoverybench

- Category: [Knowledge](/knowledge)
- Last updated: September 29, 2026 snapshot

## About DrugDiscoveryBench

- Year: 2026
- Tasks: 23 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/drugdiscoverybench)

BenchLM mirrors 23 published rows from the DrugDiscoveryBench public table captured on September 29, 2026 snapshot.

DrugDiscoveryBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (23 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT 6 Astra (mini-SWE-agent) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | OpenAI | 68.70% |
| 2 | [GPT 6 Astra (Codex) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | OpenAI | 65.40% |
| 3 | [Muse-Spark 1.3 (mini-SWE-agent) high](https://labs.scale.com/leaderboard/drugdiscoverybench) | Meta | 62.20% |
| 4 | [GPT 5.6 Sol (Codex) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | OpenAI | 56.50% |
| 5 | [Claude Opus 5 (Claude Code) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | Anthropic | 53.70% |
| 6 | [GPT-5.5 (mini-SWE-agent) xhigh](https://labs.scale.com/leaderboard/drugdiscoverybench) | OpenAI | 51.60% |
| 7 | [Claude Sonnet 5 (mini-swe-agent) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | Anthropic | 50.04% |
| 8 | [Gemini 3.5 Flash (Gemini CLI) high](https://labs.scale.com/leaderboard/drugdiscoverybench) | Google | 50.00% |
| 9 | [Gemini 3.5 Flash (mini-SWE-agent) high](https://labs.scale.com/leaderboard/drugdiscoverybench) | Google | 48.80% |
| 10 | [Claude Opus 4.8 (mini-SWE-agent) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | Anthropic | 46.80% |
| 11 | [Claude Opus 4.8 (Claude Code) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | Anthropic | 45.10% |
| 12 | [GPT-5.5 (Codex) xhigh](https://labs.scale.com/leaderboard/drugdiscoverybench) | OpenAI | 45.10% |
| 13 | [Claude Sonnet 5.0 (Claude Code) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | Anthropic | 44.70% |
| 14 | [Gemini 3.1 Pro (Gemini CLI) high](https://labs.scale.com/leaderboard/drugdiscoverybench) | OpenAI | 41.90% |
| 15 | [GLM 5.2 (mini-SWE-agent) xhigh](https://labs.scale.com/leaderboard/drugdiscoverybench) | Z.AI | 36.20% |
| 16 | [Kimi K2.7 Code (mini-SWE-agent) xhigh](https://labs.scale.com/leaderboard/drugdiscoverybench) | Moonshot AI | 35.30% |
| 17 | [DeepSeek V4 Pro (mini-SWE-agent) xhigh](https://labs.scale.com/leaderboard/drugdiscoverybench) | Deepseek | 31.70% |
| 18 | [Claude Sonnet 4.6 (mini-SWE-agent) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | Anthropic | 31.30% |
| 19 | [GPT-5.2 (Codex) xhigh](https://labs.scale.com/leaderboard/drugdiscoverybench) | OpenAI | 29.30% |
| 20 | [Qwen 3.7 Max (mini-SWE-agent) xhigh](https://labs.scale.com/leaderboard/drugdiscoverybench) | Qwen | 29.30% |
| 21 | [Claude Opus 4.6 (Claude Code) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | Anthropic | 27.70% |
| 22 | [Claude Sonnet 4.6 (Claude Code) max](https://labs.scale.com/leaderboard/drugdiscoverybench) | Anthropic | 24.00% |
| 23 | [MiniMax M3 (mini-SWE-agent) xhigh](https://labs.scale.com/leaderboard/drugdiscoverybench) | MiniMax | 22.80% |

## FAQ

### What does DrugDiscoveryBench measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published DrugDiscoveryBench snapshot?

GPT 6 Astra (mini-SWE-agent) max currently leads the published DrugDiscoveryBench snapshot with a score of 68.70%.

### How many models are evaluated on DrugDiscoveryBench?

The September 29, 2026 snapshot contains 23 AI models.

### Does DrugDiscoveryBench affect BenchLM's overall score?

Not directly. DrugDiscoveryBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
