# CVE-Bench v1 Zero-Day Black-Box Evaluation (CVE-Bench zero-day)

> OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.

Canonical page: https://benchlm.ai/benchmarks/cvebenchzerodayblackbox

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About CVE-Bench zero-day

- Year: 2026
- Tasks: 40 critical CVEs
- Format: Pass@1 over three rollouts
- Difficulty: Black-box vulnerability exploitation
- Paper: [CVE-Bench](https://github.com/uiuc-kang-lab/cve-bench)

OpenAI reports that the evaluation uses zero-day-style prompts, no source access, and pass@1 over three rollouts. BenchLM documents the protocol but stores no model score until an exact machine-readable value is available.

CVE-Bench zero-day is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### What does CVE-Bench zero-day measure?

OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.

### Which model scores highest on CVE-Bench zero-day?

No models have been evaluated on CVE-Bench zero-day yet.

### How many models are evaluated on CVE-Bench zero-day?

0 AI models have been evaluated on CVE-Bench zero-day on BenchLM.

### Does CVE-Bench zero-day affect BenchLM's overall score?

Not directly. CVE-Bench zero-day is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
