# ExploitBench v8-bench (ExploitBench)

> A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags.

Canonical page: https://benchlm.ai/benchmarks/exploitbench

- Category: [external](/external)
- Last updated: May 18, 2026

## About ExploitBench

- Year: 2026
- Tasks: V8 exploit synthesis runs
- Format: Capability coverage percentage over 16 flags
- Difficulty: Browser exploitation and cybersecurity
- Paper: [ExploitBench](https://exploitbench.ai/)

ExploitBench measures whether LLM agents can turn patched V8 bugs into progressively stronger exploit capabilities, from reaching vulnerable code to full control. BenchLM mirrors the official public leaderboard as display-only security-evaluation context.

ExploitBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (7 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Mythos Preview](/models/claude-mythos-preview) | Anthropic | 69% |
| 2 | [Claude Mythos Preview](/models/claude-mythos-preview) | Anthropic | 68% |
| 3 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 41% |
| 4 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 34% |
| 5 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 33% |
| 6 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 29% |
| 7 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 27% |

## FAQ

### What does ExploitBench measure?

A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags.

### Which model leads the published ExploitBench snapshot?

Claude Mythos Preview currently leads the published ExploitBench snapshot with a score of 69%.

### How many models are evaluated on ExploitBench?

The May 18, 2026 contains 7 AI models.

### Does ExploitBench affect BenchLM's overall score?

Not directly. ExploitBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
