CyScenarioBench Average Success Rate (CyScenarioBench success)
We show this table for reference; we do not rank on it.
Average success rate across realistic, long-horizon cybersecurity scenarios.
Benchmark score on CyScenarioBench success — September 27, 2026
We compile the CyScenarioBench success rows from benchmark-owner or independent runs. GPT-6 Astra leads the table at 59%, followed by GPT-5.6 Sol (28%). We do not use these results to rank models overall.
GPT-6 Astra
OpenAI
GPT-5.6 Sol
OpenAI
2 modelsAgenticCurrentDisplay onlyUpdated September 27, 2026
Benchmark score table (2 models)
ScoreAbout CyScenarioBench success
Year
2026
Tasks
11 cyber scenarios
Format
Average success rate
Difficulty
Long-horizon cybersecurity
Irregular evaluates models repeatedly across CyScenarioBench challenges and averages success equally across scenarios. BenchLM keeps this independent result display only.
Freshness and provenance
Version
CyScenarioBench success 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does CyScenarioBench success measure?
Average success rate across realistic, long-horizon cybersecurity scenarios.
Which model scores highest on CyScenarioBench success?
GPT-6 Astra by OpenAI currently leads with a score of 59% on CyScenarioBench success.
How many models are evaluated on CyScenarioBench success?
2 AI models have been evaluated on CyScenarioBench success on BenchLM.
Compare top models on CyScenarioBench success
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.