Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

CVE-Bench v1 Zero-Day Black-Box Evaluation (CVE-Bench zero-day)

OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.

Data verified 27 confirmed releases in the last 30 daysStart free brief

About CVE-Bench zero-day

Year

2026

Tasks

40 critical CVEs

Format

Pass@1 over three rollouts

Difficulty

Black-box vulnerability exploitation

OpenAI reports that the evaluation uses zero-day-style prompts, no source access, and pass@1 over three rollouts. BenchLM documents the protocol but stores no model score until an exact machine-readable value is available.

BenchLM freshness & provenance

Version

CVE-Bench zero-day 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does CVE-Bench zero-day measure?

OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.

Which model scores highest on CVE-Bench zero-day?

No models have been evaluated on CVE-Bench zero-day yet.

How many models are evaluated on CVE-Bench zero-day?

0 AI models have been evaluated on CVE-Bench zero-day on BenchLM.

Last updated: August 13, 2026 · BenchLM version CVE-Bench zero-day 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.