Skip to main content
BenchLM

Bug Hunt Bench

We show this table for reference; we do not rank on it.

A blind-graded coding-agent benchmark with 105 planted bugs across two production TypeScript repositories.

Planted bugs fixed on Bug Hunt Bench — September 23, 2026 snapshot

We mirror the published planted bugs fixed view for Bug Hunt Bench. Claude Fable 5.1 leads the public snapshot at 43 fixes, followed by GPT-5.6 Sol (42 fixes) and GPT-5.6 Sol (39 fixes). We do not use these results to rank models overall.

14 effort-level runsCodingCurrentDisplay onlyUpdated September 23, 2026 snapshot

Planted bugs fixed table (14 effort-level runs)

Score
1
Claude Fable 5.1Anthropic · ClosedClaude Code · max reasoning$87.18 list-rate estimate
43 fixes
2
GPT-5.6 SolOpenAI · ClosedCodex CLI · max reasoning$69.61 list-rate estimate
42 fixes
3
GPT-5.6 SolOpenAI · ClosedCodex CLI · xhigh reasoning$52.75 list-rate estimate
39 fixes
4
GPT-5.6 SolOpenAI · ClosedCodex CLI · high reasoning$33.92 list-rate estimate
34 fixes
5
Claude Fable 5.1Anthropic · ClosedClaude Code · high reasoning$48.58 list-rate estimate
33 fixes
6
Claude Fable 5.1Anthropic · ClosedClaude Code · low reasoning$33.00 list-rate estimate
29 fixes
7
Grok 4.6xAI · ClosedGrok Build CLI (ACP) · xhigh reasoning$16.96 lower bound
27 fixes
8
Claude Opus 5Anthropic · ClosedClaude Code · max reasoning$54.94 list-rate estimate
27 fixes
9
Claude Opus 5Anthropic · ClosedClaude Code · xhigh reasoning$63.25 list-rate estimate
26 fixes
10
Claude Opus 5Anthropic · ClosedClaude Code · medium reasoning$37.40 list-rate estimate
24 fixes
11
Grok 4.6xAI · ClosedGrok Build CLI (ACP) · high reasoning$15.73 lower bound
23 fixes
12
Grok 4.6xAI · ClosedGrok Build CLI (ACP) · medium reasoning$5.80 lower bound
22 fixes
13
Claude Opus 5Anthropic · ClosedClaude Code · high reasoning$41.51 list-rate estimate
21 fixes
14
Grok 4.6xAI · ClosedGrok Build CLI (ACP) · low reasoning$4.14 lower bound
15 fixes

How we show Bug Hunt Bench

The supplied view compares 14 effort-level runs across 4 models. Each run receives one round on two production TypeScript repositories containing 105 planted bugs, and an independent judge grades the submitted diff against a withheld answer key.

Only verified fixes to planted bugs enter the score. Partial fixes, claimed-only findings, and genuine unplanted fixes stay separate. Across all 208 source runs, 31 bugs remain unfixed by every model.

We keep every effort setting and agent client visible. Costs mix list-rate estimates and reconstructed lower bounds, so the table is display only and does not enter model rankings.

Snapshot

14 selected runs4 models105 planted bugs2 production repositories31 bugs still unfixedDisplay only

The published Bug Hunt Bench snapshot places Claude Fable 5.1 first at 43 fixes. The third row is 4 points behind. The broader top-10 range is 19 points, so the table still separates the published systems.

14 effort-level runs have been evaluated on Bug Hunt Bench. The benchmark falls in the Coding category. Bug Hunt Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Bug Hunt Bench

Year

2026

Tasks

105 planted bugs across two production TypeScript repositories

Format

Strict planted bugs fixed

Difficulty

Blind production-repository bug finding and repair

Each run gives a model the same prompt and one round per repository in its native agentic CLI. An independent judge checks the submitted diff against a withheld answer key. The score counts only planted bugs fixed; partial fixes, claimed-only findings, and genuine unplanted fixes remain separate. BenchLM mirrors the supplied 14-run effort comparison as display-only system evidence.

Freshness and provenance

Version

Bug Hunt Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Bug Hunt Bench measure?

Bug Hunt Bench measures whether coding agents can find and fix 105 planted bugs across two production TypeScript repositories. Every run gets the same prompt and one attempt per repository. An independent judge grades the submitted diff against a withheld answer key rather than trusting the model's report.

Which run leads the selected Bug Hunt Bench view?

Claude Fable 5.1 at max effort leads the supplied September 2, 2026 selection with 43 of 105 planted bugs fixed. GPT-5.6 Sol at max effort follows with 42. These are complete model, reasoning-effort, and CLI configurations, not isolated base-model measurements.

Why is Bug Hunt Bench display only?

Each row combines the model with a reasoning setting, agentic CLI, one run per repository, and a cost figure that may be a list-rate estimate or reconstructed lower bound. The snapshot is useful for comparing those exact setups, but it does not enter BenchLM's model-only rankings.

Last updated: September 23, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.