Benchmark profile
Weapons of Mass Destruction Proxy — Chemistry (WMDP Chem)
The chemical hazardous-knowledge split of the public WMDP benchmark.
Data verified 26 confirmed releases in the last 30 daysStart free briefBenchmark score on WMDP Chem — August 12, 2026
BenchLM mirrors the published score view for WMDP Chem. Grok 4.5 leads the public snapshot at 87.3% , followed by Grok 4.6 (85.3%). We do not use these results to rank models overall.
Grok 4.5
xAI
grok-4-5
Grok 4.6
xAI
grok-4-6
About WMDP Chem
Year
2024
Tasks
Hazardous chemical-knowledge multiple-choice questions
Format
Accuracy
Difficulty
Chemical-safety knowledge
WMDP releases its data and evaluation code publicly. BenchLM keeps model-card comparison rows display-only because protocol details can vary across providers.
BenchLM freshness & provenance
Version
WMDP Chem 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does WMDP Chem measure?
The chemical hazardous-knowledge split of the public WMDP benchmark.
Which model scores highest on WMDP Chem?
Grok 4.5 by xAI currently leads with a score of 87.3% on WMDP Chem.
How many models are evaluated on WMDP Chem?
2 AI models have been evaluated on WMDP Chem on BenchLM.
Compare Top Models on WMDP Chem
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.