# SRE-Bench binary reverse engineering (SRE-Bench)

> A contamination-controlled benchmark that asks agents to reverse engineer compiled binaries without source access and satisfy six task-specific objectives per challenge.

Canonical page: https://benchlm.ai/benchmarks/srebench

- Category: [external](/external)
- Last updated: September 10, 2026

## About SRE-Bench

- Year: 2026
- Tasks: 262 binary reverse-engineering instances
- Format: Fully solved challenge rate (pass@1)
- Difficulty: Binary reverse engineering
- Paper: [GPT-6 Astra System Card](https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf)

SRE-Bench contains 262 binary instances derived from 19 privately developed programs spanning network protocols, file formats, malware remediation, and firmware across C, C++, Go, and Rust. A challenge counts as solved only when all six objectives are satisfied. BenchLM stores the single-attempt launch-table value as a display-only external security row and records pass@4 results in provenance notes.

SRE-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 88.0% |

## FAQ

### What does SRE-Bench measure?

A contamination-controlled benchmark that asks agents to reverse engineer compiled binaries without source access and satisfy six task-specific objectives per challenge.

### Which model scores highest on SRE-Bench?

GPT-6 Astra by OpenAI currently leads with a score of 88.0% on SRE-Bench.

### How many models are evaluated on SRE-Bench?

1 AI models have been evaluated on SRE-Bench on BenchLM.

### Does SRE-Bench affect BenchLM's overall score?

Not directly. SRE-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
