# ExploitGym

> A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.

Canonical page: https://benchlm.ai/benchmarks/exploitgym

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About ExploitGym

- Year: 2026
- Tasks: 898 exploitation tasks
- Format: Working exploit generation
- Difficulty: Advanced cybersecurity exploitation
- Paper: [ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?](https://arxiv.org/abs/2605.11086)

ExploitGym contains 898 containerized tasks sourced from real-world vulnerabilities across userspace programs, V8, and the Linux kernel. BenchLM normalizes successful exploit counts to percentage of the 898-task suite and keeps the benchmark display-only because it measures dual-use offensive cybersecurity capability.

ExploitGym is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (10 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 42.4% |
| 2 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 33.7% |
| 3 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 23.2% |
| 4 | [Claude Mythos Preview](/models/claude-mythos-preview) | Anthropic | 17.5% |
| 5 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 15.3% |
| 6 | [GLM-5.3](/models/glm-5-3) | Z.AI | 15.0% |
| 7 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 13.4% |
| 8 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 12.4% |
| 9 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 6.0% |
| 10 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 0.8% |

## FAQ

### What does ExploitGym measure?

A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.

### Which model scores highest on ExploitGym?

GPT-6 Astra by OpenAI currently leads with a score of 42.4% on ExploitGym.

### How many models are evaluated on ExploitGym?

10 AI models have been evaluated on ExploitGym on BenchLM.

### Does ExploitGym affect BenchLM's overall score?

Not directly. ExploitGym is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on ExploitGym

- [GPT-6 Astra vs GPT-5.6 Sol](/compare/gpt-5-6-sol-vs-gpt-6-astra)
- [GPT-5.6 Sol vs GPT-5.6 Terra](/compare/gpt-5-6-sol-vs-gpt-5-6-terra)
- [GPT-5.6 Terra vs Claude Mythos Preview](/compare/claude-mythos-preview-vs-gpt-5-6-terra)
- [Claude Mythos Preview vs DeepSeek V4.1 Flash](/compare/claude-mythos-preview-vs-deepseek-v4-1-flash)
