Benchmark profile
Bug Hunt Bench: Proactive AI Bug Discovery
A proactive bug-fixing benchmark that asks coding agents to find and repair hidden defects across two working TypeScript applications without an issue report or failing test pointing to the answers.
The public Bug Hunt Bench best-run snapshot snapshot ranks GPT-5.6 Sol first at 42/105, ahead of GPT-5.6 Luna (33/105) and Fable 5 (29/105) among 13 tested models. We mirror the table as display-only evidence; it does not affect overall rankings.
How to read this leaderboard
Higher is better. A point means the submitted diff fully fixed one hidden bug; partial fixes and report-only claims receive no credit. This page shows each named model release's best published run and keeps effort level and native harness visible beside the score.
Operator receipt: 13 sourced rows are currently displayable on this page; the leading published row is GPT-5.6 Sol at 42/105.
Honest limit: Each cell has one run per repository, both repositories use TypeScript, and models run in different native harnesses. The best-run view favors models tested at more settings. Answer keys and model diffs are withheld, while an LLM judge supplies the published verdicts after cross-vendor calibration.
Can an agent fix bugs nobody reported?
We mirror Bug Hunt Bench's August 12, 2026 snapshot: 105 withheld bugs across two working TypeScript applications whose checks still pass. 76 are real shipped regressions and 29 were written in the first repository's style. Each agent gets the same open brief to find and fix as many bugs as it can without an issue report or a failing test pointing to the answer.
The visible score is the best published strict-fix run for each of 13 named model releases. A blind judge grades the submitted diff against the withheld answer key, gives no partial credit, and ignores claims that do not appear in the patch. The snapshot preserves all 22 source arms, including lower-effort runs and reruns, so the selection rule remains auditable.
Bug Hunt Bench stays display only. It uses one run per repository and setting, two TypeScript codebases, native vendor harnesses instead of one controlled agent scaffold, and withheld answer keys and diffs. Selecting each model’s best published run also favors models tested at more settings. The result is useful product-level evidence, not a normalized base-model score.
Strict fixes out of 105 on Bug Hunt Bench best-run snapshot — August 12, 2026 snapshot
BenchLM mirrors the published strict fixes out of 105 view for Bug Hunt Bench best-run snapshot. GPT-5.6 Sol leads the public snapshot at 42/105 , followed by GPT-5.6 Luna (33/105) and Fable 5 (29/105). We do not use these results to rank models overall.
GPT-5.6 Sol
OpenAI
max · Codex CLI
GPT-5.6 Sol (max effort)
GPT-5.6 Luna
OpenAI
max · Codex CLI
GPT-5.6 Luna (max effort)
Fable 5
Anthropic
max · Claude Code
Fable 5 (max effort)
Strict fixes out of 105 table (13 models)
ScoreThe published Bug Hunt Bench snapshot places GPT-5.6 Sol first at 42/105. The third row is 13 points behind. The broader top-10 range is 28 points, so the table still separates the published systems.
13 models have been evaluated on Bug Hunt Bench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Bug Hunt Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Bug Hunt Bench
Year
2026
Tasks
105 hidden bugs across 2 TypeScript applications
Format
Strict fixes from one run per repository and setting
Difficulty
Proactive repository-wide bug discovery and repair
Bug Hunt Bench hides 105 defects inside two working TypeScript applications while preserving their existing green checks. Seventy-six defects are regressions that shipped and were later fixed; 29 were authored in the style of the first repository. Agents receive no issue reports, use their native coding harnesses, edit the repositories in place, and are graded from their diffs against withheld answer keys by a blind cross-family judge.
BenchLM freshness & provenance
Version
bugHuntBench
Refresh cadence
Static
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Bug Hunt Bench measure?
Bug Hunt Bench measures whether an AI coding agent can discover and repair bugs without an issue report or failing test. Agents inspect two working TypeScript repositories with 105 hidden defects, edit source in place, and keep existing checks green. A blind judge counts only complete fixes present in the submitted diff.
How is Bug Hunt Bench different from SWE-bench?
SWE-bench gives an agent a specific GitHub issue and asks for one patch. Bug Hunt Bench removes that guidance and hides many defects inside an otherwise working repository. The agent must decide where to look, identify multiple bugs, repair them, and avoid breaking checks that already pass.
Which AI coding model leads Bug Hunt Bench?
GPT-5.6 Sol leads the August 12, 2026 snapshot with 42 strict fixes out of 105 at max effort in Codex CLI. That is the best published run, not an average. The page remains display only because effort, harness, reruns, and repository coverage all affect the result.
Are Bug Hunt Bench scores model-only results?
No. Every row measures a model inside a coding-agent product and serving route. GPT runs in Codex CLI, Claude-family models run in Claude Code, and Grok runs in Grok Build CLI. Harness behavior, effort support, routing, and tool use are part of the observed result.
Does Bug Hunt Bench measure security vulnerability discovery?
Not specifically. The hidden set covers general product defects rather than a security taxonomy, exploitability test, or CVE corpus. Some agents found genuine unplanted issues with security consequences, but those extras are supporting evidence. The headline score counts complete fixes to the 105 withheld benchmark bugs.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.