Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

SRE-Bench binary reverse engineering (SRE-Bench)

A contamination-controlled benchmark that asks agents to reverse engineer compiled binaries without source access and satisfy six task-specific objectives per challenge.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on SRE-Bench — September 10, 2026

We mirror the published score view for SRE-Bench. GPT-6 Astra leads the public snapshot at 88.0%. We do not use these results to rank models overall.

1 modelExternal benchmark mirrorsCurrentDisplay onlyUpdated September 10, 2026

Benchmark score table (1 model)

Score
1
GPT-6 AstraOpenAI · Closed
88.0%

About SRE-Bench

Year

2026

Tasks

262 binary reverse-engineering instances

Format

Fully solved challenge rate (pass@1)

Difficulty

Binary reverse engineering

SRE-Bench contains 262 binary instances derived from 19 privately developed programs spanning network protocols, file formats, malware remediation, and firmware across C, C++, Go, and Rust. A challenge counts as solved only when all six objectives are satisfied. BenchLM stores the single-attempt launch-table value as a display-only external security row and records pass@4 results in provenance notes.

BenchLM freshness & provenance

Version

SRE-Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SRE-Bench measure?

A contamination-controlled benchmark that asks agents to reverse engineer compiled binaries without source access and satisfy six task-specific objectives per challenge.

Which model scores highest on SRE-Bench?

GPT-6 Astra by OpenAI currently leads with a score of 88.0% on SRE-Bench.

How many models are evaluated on SRE-Bench?

1 AI models have been evaluated on SRE-Bench on BenchLM.

Last updated: September 10, 2026 · BenchLM version SRE-Bench 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.