Cybersecurity benchmarks
Best AI models for cybersecurity
There is no honest single cyber ranking yet. GPT-5.6 Sol has the broadest public coverage in this snapshot; Claude Mythos 5 leads the listed provider-run ExploitBench rows. BenchLM keeps every cyber result display only because the protocols measure different capabilities, use different harnesses, and sometimes test safeguards rather than skill.
Use this page to compare like-for-like results, inspect the source behind each number, and separate model capability from trusted-access policy.
Broadest coverage
GPT-5.6 Sol
Public rows span CyberGym, ExploitGym, ExploitBench, SEC-Bench Pro, The Last Ones, FrontierCyber, CyScenarioBench, and Atomic.
ExploitBench
Claude Mythos 5: 78%
Anthropic reports this safeguards-off result. Fable 5 does not inherit it: Anthropic published no Fable cyber result.
Trusted access
Blue is access; Red is a model
Daybreak Blue changes safeguards around GPT-5.6 Sol. Daybreak Red exposes the separate GPT-5.6 Cyber model.
Cyber benchmark comparison
Dashes mean no exact public result is attached to that model and protocol.
| Model | CyberGym | ExploitGym | ExploitBench | SEC-Bench Pro | The Last Ones | FrontierCyber |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | 84.5% | 33.7% | 73.5% | 71.2% | 70% | 9.6% |
| Claude Mythos 5 Anthropic | 83.8% | — | 78% | — | 60% | — |
| Claude Mythos Preview Anthropic | 83.1% | 17.5% | 69% | — | — | — |
| GPT-5.5 OpenAI | 81.8% | 13.4% | — | 45.8% | — | — |
| GPT-5.6 Cyber OpenAI | — | — | — | — | — | — |
These columns are not averaged. Provider-run, independent, private, long-horizon, and access-compliance results remain separate.
Cybersecurity benchmark directory
Open a benchmark profile for its protocol, source, model rows, and verification status.
CyberGym
A cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.
Vulnerability reproduction and PoC generation · display only
ExploitGym
A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.
Working exploit generation · display only
ExploitBench
A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags.
Capability coverage percentage over 16 flags · display only
SEC-Bench Pro
Cybersecurity benchmark for agentic vulnerability analysis and exploit-oriented security tasks.
Success rate · display only
FrontierCyber
Independent evaluation of AI agents against vulnerable real-world systems in dynamic environments.
Tasks solved · display only
CyScenarioBench success
Average success rate across realistic, long-horizon cybersecurity scenarios.
Average success rate · display only
Atomic vulnerability research
Irregular's domain-level evaluation of vulnerability research and exploitation capability.
Domain average · display only
SCONE post-cutoff success
Share of SCONE smart-contract vulnerabilities exploited on the 12-task post-cutoff set.
Best@8 exploit success rate · display only
Firefox 147 exploits
Share of patched Firefox 147 JavaScript-engine targets for which the model produced a working arbitrary-code-execution exploit.
Working arbitrary-code-execution rate · display only
Anthropic OSS-Fuzz crash
Share of evaluated OSS-Fuzz entry points where the model produced at least a crash.
Any-crash rate · display only
The Last Ones completion
Share of runs that completed the 32-step cyber range within the 100-million-token limit.
Completion rate · display only
ACCR Daybreak Red
Share of approved advanced-cyber requests completed rather than refused by GPT-5.6 Cyber under Daybreak Red access.
Completion rate · display only
CVE-Bench zero-day
OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.
Pass@1 over three rollouts · display only
How to read the results
Match protocols. Compare CyberGym with CyberGym and ExploitBench with ExploitBench. Do not average an exploit success rate with a cyber-range step count.
Check the harness. Agent scaffolds, time limits, token budgets, source-code access, and Best@k settings can move the result as much as the model.
Separate access from capability. Advanced Cyber Completion Rate measures whether approved requests are completed rather than refused. A high ACCR is not a high cyber skill score.
Treat private evaluations cautiously. Provider system cards add useful evidence, but independent and benchmark-native replications carry a different evidentiary weight.
Daybreak access map
- Standard
- GPT-5.6 Sol with standard safeguards.
- Daybreak Blue
- The same GPT-5.6 Sol model through the
gpt-daybreak-bluealias, with modified safeguards for approved users. - Daybreak Red
- The separate GPT-5.6 Cyber model through the
gpt-daybreak-redalias.
Questions about AI cybersecurity benchmarks
What is the best AI model for cybersecurity?
There is no defensible single winner across the public evidence. GPT-5.6 Sol has the broadest mix of provider and independent cyber results in this snapshot, while Claude Mythos 5 leads the listed provider-run ExploitBench comparison. Access, safeguards, harnesses, and task types differ, so compare models within the same benchmark protocol.
Are Daybreak Blue and Daybreak Red separate models?
Daybreak Blue is an access configuration for GPT-5.6 Sol with modified cyber safeguards. Daybreak Red provides trusted access to the separate GPT-5.6 Cyber model.
Why is there no combined cybersecurity score?
CyberGym, ExploitGym, ExploitBench, FrontierCyber, SCONE, and ACCR measure different things under different harnesses. Averaging them would hide protocol differences and could turn an access-refusal metric into a capability claim.