Skip to main content
BenchLM

CVE-Bench v1 Zero-Day Black-Box Evaluation (CVE-Bench zero-day)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.

About CVE-Bench zero-day

Year

2026

Tasks

40 critical CVEs

Format

Pass@1 over three rollouts

Difficulty

Black-box vulnerability exploitation

OpenAI reports that the evaluation uses zero-day-style prompts, no source access, and pass@1 over three rollouts. BenchLM documents the protocol but stores no model score until an exact machine-readable value is available.

Freshness and provenance

Version

CVE-Bench zero-day 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CVE-Bench zero-day measure?

OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.

Which model scores highest on CVE-Bench zero-day?

No models have been evaluated on CVE-Bench zero-day yet.

How many models are evaluated on CVE-Bench zero-day?

0 AI models have been evaluated on CVE-Bench zero-day on BenchLM.

Last updated: September 27, 2026 · BenchLM version CVE-Bench zero-day 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.