CVE-Bench v1 Zero-Day Black-Box Evaluation (CVE-Bench zero-day)
OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.
About CVE-Bench zero-day
Year
2026
Tasks
40 critical CVEs
Format
Pass@1 over three rollouts
Difficulty
Black-box vulnerability exploitation
OpenAI reports that the evaluation uses zero-day-style prompts, no source access, and pass@1 over three rollouts. BenchLM documents the protocol but stores no model score until an exact machine-readable value is available.
BenchLM freshness & provenance
Version
CVE-Bench zero-day 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does CVE-Bench zero-day measure?
OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.
Which model scores highest on CVE-Bench zero-day?
No models have been evaluated on CVE-Bench zero-day yet.
How many models are evaluated on CVE-Bench zero-day?
0 AI models have been evaluated on CVE-Bench zero-day on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.