# ProgramBench hidden-test pass rate after episode 1 (ProgramBench (episode 1))

> Program-reconstruction hidden-test pass rate after the first of five sequential long-context episodes.

Canonical page: https://benchlm.ai/benchmarks/programbenchepisode1

- Category: [Coding](/coding)
- Last updated: September 27, 2026

## About ProgramBench (episode 1)

- Year: 2026
- Tasks: 166 golden program-reconstruction tasks
- Format: Hidden-test pass rate after episode 1
- Difficulty: Long-context clean-room software engineering
- Paper: [ProgramBench: Can language models rebuild programs from scratch?](https://arxiv.org/abs/2605.03546)

Anthropic evaluated 166 golden tasks with no internet or decompilation tools and a fresh context budget of up to one million tokens per episode. This is kept separate from the episode-5 result.

ProgramBench (episode 1) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 83.0% |

## FAQ

### What does ProgramBench (episode 1) measure?

Program-reconstruction hidden-test pass rate after the first of five sequential long-context episodes.

### Which model scores highest on ProgramBench (episode 1)?

Claude Opus 5 by Anthropic currently leads with a score of 83.0% on ProgramBench (episode 1).

### How many models are evaluated on ProgramBench (episode 1)?

1 AI models have been evaluated on ProgramBench (episode 1) on BenchLM.

### Does ProgramBench (episode 1) affect BenchLM's overall score?

Not directly. ProgramBench (episode 1) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
