Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Cybersecurity benchmarks

Best AI models for cybersecurity

There is no honest single cyber ranking yet. GPT-5.6 Sol has the broadest public coverage in this snapshot; Claude Mythos 5 leads the listed provider-run ExploitBench rows. BenchLM keeps every cyber result display only because the protocols measure different capabilities, use different harnesses, and sometimes test safeguards rather than skill.

Use this page to compare like-for-like results, inspect the source behind each number, and separate model capability from trusted-access policy.

Broadest coverage

GPT-5.6 Sol

Public rows span CyberGym, ExploitGym, ExploitBench, SEC-Bench Pro, The Last Ones, FrontierCyber, CyScenarioBench, and Atomic.

ExploitBench

Claude Mythos 5: 78%

Anthropic reports this safeguards-off result. Fable 5 does not inherit it: Anthropic published no Fable cyber result.

Trusted access

Blue is access; Red is a model

Daybreak Blue changes safeguards around GPT-5.6 Sol. Daybreak Red exposes the separate GPT-5.6 Cyber model.

Cyber benchmark comparison

Dashes mean no exact public result is attached to that model and protocol.

ModelCyberGymExploitGymExploitBenchSEC-Bench ProThe Last OnesFrontierCyber
GPT-5.6 Sol

OpenAI

84.5%33.7%73.5%71.2%70%9.6%
Claude Mythos 5

Anthropic

83.8%78%60%
Claude Mythos Preview

Anthropic

83.1%17.5%69%
GPT-5.5

OpenAI

81.8%13.4%45.8%
GPT-5.6 Cyber

OpenAI

These columns are not averaged. Provider-run, independent, private, long-horizon, and access-compliance results remain separate.

Cybersecurity benchmark directory

Open a benchmark profile for its protocol, source, model rows, and verification status.

CyberGym

A cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.

Vulnerability reproduction and PoC generation · display only

ExploitGym

A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.

Working exploit generation · display only

ExploitBench

A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags.

Capability coverage percentage over 16 flags · display only

SEC-Bench Pro

Cybersecurity benchmark for agentic vulnerability analysis and exploit-oriented security tasks.

Success rate · display only

FrontierCyber

Independent evaluation of AI agents against vulnerable real-world systems in dynamic environments.

Tasks solved · display only

CyScenarioBench success

Average success rate across realistic, long-horizon cybersecurity scenarios.

Average success rate · display only

Atomic vulnerability research

Irregular's domain-level evaluation of vulnerability research and exploitation capability.

Domain average · display only

SCONE post-cutoff success

Share of SCONE smart-contract vulnerabilities exploited on the 12-task post-cutoff set.

Best@8 exploit success rate · display only

Firefox 147 exploits

Share of patched Firefox 147 JavaScript-engine targets for which the model produced a working arbitrary-code-execution exploit.

Working arbitrary-code-execution rate · display only

Anthropic OSS-Fuzz crash

Share of evaluated OSS-Fuzz entry points where the model produced at least a crash.

Any-crash rate · display only

The Last Ones completion

Share of runs that completed the 32-step cyber range within the 100-million-token limit.

Completion rate · display only

ACCR Daybreak Red

Share of approved advanced-cyber requests completed rather than refused by GPT-5.6 Cyber under Daybreak Red access.

Completion rate · display only

CVE-Bench zero-day

OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.

Pass@1 over three rollouts · display only

How to read the results

Match protocols. Compare CyberGym with CyberGym and ExploitBench with ExploitBench. Do not average an exploit success rate with a cyber-range step count.

Check the harness. Agent scaffolds, time limits, token budgets, source-code access, and Best@k settings can move the result as much as the model.

Separate access from capability. Advanced Cyber Completion Rate measures whether approved requests are completed rather than refused. A high ACCR is not a high cyber skill score.

Treat private evaluations cautiously. Provider system cards add useful evidence, but independent and benchmark-native replications carry a different evidentiary weight.

Daybreak access map

Standard
GPT-5.6 Sol with standard safeguards.
Daybreak Blue
The same GPT-5.6 Sol model through the gpt-daybreak-blue alias, with modified safeguards for approved users.
Daybreak Red
The separate GPT-5.6 Cyber model through the gpt-daybreak-red alias.

Questions about AI cybersecurity benchmarks

What is the best AI model for cybersecurity?

There is no defensible single winner across the public evidence. GPT-5.6 Sol has the broadest mix of provider and independent cyber results in this snapshot, while Claude Mythos 5 leads the listed provider-run ExploitBench comparison. Access, safeguards, harnesses, and task types differ, so compare models within the same benchmark protocol.

Are Daybreak Blue and Daybreak Red separate models?

Daybreak Blue is an access configuration for GPT-5.6 Sol with modified cyber safeguards. Daybreak Red provides trusted access to the separate GPT-5.6 Cyber model.

Why is there no combined cybersecurity score?

CyberGym, ExploitGym, ExploitBench, FrontierCyber, SCONE, and ACCR measure different things under different harnesses. Averaging them would hide protocol differences and could turn an access-refusal metric into a capability claim.