Bug Hunt Bench
We show this table for reference; we do not rank on it.
A blind-graded coding-agent benchmark with 105 planted bugs across two production TypeScript repositories.
Planted bugs fixed on Bug Hunt Bench — September 23, 2026 snapshot
We mirror the published planted bugs fixed view for Bug Hunt Bench. Claude Fable 5.1 leads the public snapshot at 43 fixes, followed by GPT-5.6 Sol (42 fixes) and GPT-5.6 Sol (39 fixes). We do not use these results to rank models overall.
Claude Fable 5.1
Anthropic
Claude Code · max reasoning
$87.18 list-rate estimate
GPT-5.6 Sol
OpenAI
Codex CLI · max reasoning
$69.61 list-rate estimate
GPT-5.6 Sol
OpenAI
Codex CLI · xhigh reasoning
$52.75 list-rate estimate
14 effort-level runsCodingCurrentDisplay onlyUpdated September 23, 2026 snapshot
Planted bugs fixed table (14 effort-level runs)
ScoreHow we show Bug Hunt Bench
The supplied view compares 14 effort-level runs across 4 models. Each run receives one round on two production TypeScript repositories containing 105 planted bugs, and an independent judge grades the submitted diff against a withheld answer key.
Only verified fixes to planted bugs enter the score. Partial fixes, claimed-only findings, and genuine unplanted fixes stay separate. Across all 208 source runs, 31 bugs remain unfixed by every model.
We keep every effort setting and agent client visible. Costs mix list-rate estimates and reconstructed lower bounds, so the table is display only and does not enter model rankings.
Snapshot
The published Bug Hunt Bench snapshot places Claude Fable 5.1 first at 43 fixes. The third row is 4 points behind. The broader top-10 range is 19 points, so the table still separates the published systems.
14 effort-level runs have been evaluated on Bug Hunt Bench. The benchmark falls in the Coding category. Bug Hunt Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Bug Hunt Bench
Year
2026
Tasks
105 planted bugs across two production TypeScript repositories
Format
Strict planted bugs fixed
Difficulty
Blind production-repository bug finding and repair
Each run gives a model the same prompt and one round per repository in its native agentic CLI. An independent judge checks the submitted diff against a withheld answer key. The score counts only planted bugs fixed; partial fixes, claimed-only findings, and genuine unplanted fixes remain separate. BenchLM mirrors the supplied 14-run effort comparison as display-only system evidence.
Freshness and provenance
Version
Bug Hunt Bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Bug Hunt Bench measure?
Bug Hunt Bench measures whether coding agents can find and fix 105 planted bugs across two production TypeScript repositories. Every run gets the same prompt and one attempt per repository. An independent judge grades the submitted diff against a withheld answer key rather than trusting the model's report.
Which run leads the selected Bug Hunt Bench view?
Claude Fable 5.1 at max effort leads the supplied September 2, 2026 selection with 43 of 105 planted bugs fixed. GPT-5.6 Sol at max effort follows with 42. These are complete model, reasoning-effort, and CLI configurations, not isolated base-model measurements.
Why is Bug Hunt Bench display only?
Each row combines the model with a reasoning setting, agentic CLI, one run per repository, and a cost figure that may be a list-rate estimate or reconstructed lower bound. The snapshot is useful for comparing those exact setups, but it does not enter BenchLM's model-only rankings.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.