Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Benchmark profile

Bug Hunt Bench: Proactive AI Bug Discovery

A proactive bug-fixing benchmark that asks coding agents to find and repair hidden defects across two working TypeScript applications without an issue report or failing test pointing to the answers.

The public Bug Hunt Bench best-run snapshot snapshot ranks GPT-5.6 Sol first at 42/105, ahead of GPT-5.6 Luna (33/105) and Fable 5 (29/105) among 13 tested models. We mirror the table as display-only evidence; it does not affect overall rankings.

How to read this leaderboard

Higher is better. A point means the submitted diff fully fixed one hidden bug; partial fixes and report-only claims receive no credit. This page shows each named model release's best published run and keeps effort level and native harness visible beside the score.

Operator receipt: 13 sourced rows are currently displayable on this page; the leading published row is GPT-5.6 Sol at 42/105.

Honest limit: Each cell has one run per repository, both repositories use TypeScript, and models run in different native harnesses. The best-run view favors models tested at more settings. Answer keys and model diffs are withheld, while an LLM judge supplies the published verdicts after cross-vendor calibration.

Can an agent fix bugs nobody reported?

We mirror Bug Hunt Bench's August 12, 2026 snapshot: 105 withheld bugs across two working TypeScript applications whose checks still pass. 76 are real shipped regressions and 29 were written in the first repository's style. Each agent gets the same open brief to find and fix as many bugs as it can without an issue report or a failing test pointing to the answer.

The visible score is the best published strict-fix run for each of 13 named model releases. A blind judge grades the submitted diff against the withheld answer key, gives no partial credit, and ignores claims that do not appear in the patch. The snapshot preserves all 22 source arms, including lower-effort runs and reruns, so the selection rule remains auditable.

Bug Hunt Bench stays display only. It uses one run per repository and setting, two TypeScript codebases, native vendor harnesses instead of one controlled agent scaffold, and withheld answer keys and diffs. Selecting each model’s best published run also favors models tested at more settings. The result is useful product-level evidence, not a normalized base-model score.

13 model releases22 published arms105 hidden bugs76 shipped regressions2 TypeScript repositoriesDisplay only

Strict fixes out of 105 on Bug Hunt Bench best-run snapshot — August 12, 2026 snapshot

BenchLM mirrors the published strict fixes out of 105 view for Bug Hunt Bench best-run snapshot. GPT-5.6 Sol leads the public snapshot at 42/105 , followed by GPT-5.6 Luna (33/105) and Fable 5 (29/105). We do not use these results to rank models overall.

13 modelsCodingRefreshingDisplay onlyUpdated August 12, 2026 snapshot

Strict fixes out of 105 table (13 models)

Score
1
GPT-5.6 SolOpenAI · Closedmax · Codex CLI
42/105
2
GPT-5.6 LunaOpenAI · Closedmax · Codex CLI
33/105
3
Fable 5Anthropic · Closedmax · Claude Code
29/105
4
Opus 5Anthropicmax · Claude Code
27/105
5
Grok 4.6xAI · Closedxhigh · Grok Build CLI
27/105
6
Kimi K3Moonshot AI · Closeddefault · Claude Code / OpenRouter
21/105
7
Qwen3.8-MaxAlibaba · Closedxhigh · Claude Code / Alibaba API
19/105
8
Grok 4.5xAI · Closedhigh · Grok Build CLI
17/105
9
Muse Spark 1.2Metaxhigh · Claude Code / Meta API
17/105
10
DeepSeek V4-Flash 0731DeepSeek · Closeddefault · Claude Code / OpenRouter
14/105
11
DeepSeek V4-ProDeepSeek · Closeddefault · Claude Code / OpenRouter
10/105
12
Opus 4.8Anthropic · Closedhigh · Claude Code
9/105
13
Sonnet 5Anthropic · Closedhigh · Claude Code
9/105

The published Bug Hunt Bench snapshot places GPT-5.6 Sol first at 42/105. The third row is 13 points behind. The broader top-10 range is 28 points, so the table still separates the published systems.

13 models have been evaluated on Bug Hunt Bench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Bug Hunt Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Bug Hunt Bench

Year

2026

Tasks

105 hidden bugs across 2 TypeScript applications

Format

Strict fixes from one run per repository and setting

Difficulty

Proactive repository-wide bug discovery and repair

Bug Hunt Bench hides 105 defects inside two working TypeScript applications while preserving their existing green checks. Seventy-six defects are regressions that shipped and were later fixed; 29 were authored in the style of the first repository. Agents receive no issue reports, use their native coding harnesses, edit the repositories in place, and are graded from their diffs against withheld answer keys by a blind cross-family judge.

BenchLM freshness & provenance

Version

bugHuntBench

Refresh cadence

Static

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Bug Hunt Bench measure?

Bug Hunt Bench measures whether an AI coding agent can discover and repair bugs without an issue report or failing test. Agents inspect two working TypeScript repositories with 105 hidden defects, edit source in place, and keep existing checks green. A blind judge counts only complete fixes present in the submitted diff.

How is Bug Hunt Bench different from SWE-bench?

SWE-bench gives an agent a specific GitHub issue and asks for one patch. Bug Hunt Bench removes that guidance and hides many defects inside an otherwise working repository. The agent must decide where to look, identify multiple bugs, repair them, and avoid breaking checks that already pass.

Which AI coding model leads Bug Hunt Bench?

GPT-5.6 Sol leads the August 12, 2026 snapshot with 42 strict fixes out of 105 at max effort in Codex CLI. That is the best published run, not an average. The page remains display only because effort, harness, reruns, and repository coverage all affect the result.

Are Bug Hunt Bench scores model-only results?

No. Every row measures a model inside a coding-agent product and serving route. GPT runs in Codex CLI, Claude-family models run in Claude Code, and Grok runs in Grok Build CLI. Harness behavior, effort support, routing, and tool use are part of the observed result.

Does Bug Hunt Bench measure security vulnerability discovery?

Not specifically. The hidden set covers general product defects rather than a security taxonomy, exploitability test, or CVE corpus. Some agents found genuine unplanted issues with security consequences, but those extras are supporting evidence. The headline score counts complete fixes to the 105 withheld benchmark bugs.

Last updated: August 12, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.