Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

SEC-Bench Pro

Cybersecurity benchmark for agentic vulnerability analysis and exploit-oriented security tasks.

Data verified 27 confirmed releases in the last 30 daysStart free brief

Benchmark score on SEC-Bench Pro — August 13, 2026

BenchLM mirrors the published score view for SEC-Bench Pro. GPT-5.6 Sol leads the public snapshot at 71.2% , followed by GPT-5.5 (45.8%). We do not use these results to rank models overall.

2 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated August 13, 2026

Benchmark score table (2 models)

Score
1
GPT-5.6 SolOpenAI · Closed
71.2%
2
GPT-5.5OpenAI · Closed
45.8%

About SEC-Bench Pro

Year

2026

Tasks

Security engineering tasks

Format

Success rate

Difficulty

Advanced cybersecurity

OpenAI reports exact model results on SEC-Bench Pro in its GPT-5.6 launch table. BenchLM stores the provider-run values as display-only cyber evidence until benchmark-native results are available.

BenchLM freshness & provenance

Version

SEC-Bench Pro 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SEC-Bench Pro measure?

Cybersecurity benchmark for agentic vulnerability analysis and exploit-oriented security tasks.

Which model scores highest on SEC-Bench Pro?

GPT-5.6 Sol by OpenAI currently leads with a score of 71.2% on SEC-Bench Pro.

How many models are evaluated on SEC-Bench Pro?

2 AI models have been evaluated on SEC-Bench Pro on BenchLM.

Compare Top Models on SEC-Bench Pro

Last updated: August 13, 2026 · BenchLM version SEC-Bench Pro 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.