Skip to main content
BenchLM

KernelBench Hard H100 (KernelBench)

We show this table for reference; we do not rank on it.

An agentic GPU-kernel benchmark that measures how much of the hardware roofline a model's correct, audit-clean kernels reach on six demanding CUDA and Triton problems.

Mean peak fraction of roofline on KernelBench — September 23, 2026 snapshot

We mirror the published mean peak fraction of roofline view for KernelBench. Claude Fable 5 leads the public snapshot at 23.5%, followed by Claude Opus 5 (21.8%) and Kimi K3 (20.9%). We do not use these results to rank models overall.

5 modelsCodingCurrentDisplay onlyUpdated September 23, 2026 snapshot

Mean peak fraction of roofline table (5 models)

Score
1
Claude Fable 5Anthropic · Closedclaude · max
23.5%
2
Claude Opus 5Anthropic · Closedor-opus · max
21.8%
3
Kimi K3Moonshot AI · Closedkinetic-claude
20.9%
4
GPT-5.6 SolOpenAI · Closedcodex · xhigh
13.6%
5
Qwen3.8 MaxAlibaba · Open weightor-fable · xhigh
13.4%

How we show KernelBench

We mirror the KernelBench Hard hard board from kernelbench.com. The visible score is per-problem · rank by passes · bar = share of best on that problem across 6 GPU-kernel problems; failed and audit-flagged cells do not enter that mean.

This is an independent agent benchmark, not the Stanford KernelBench project leaderboard. Rows combine a model, an agent harness, long-running GPU sessions, and human audit decisions, so the table stays display only.

Snapshot

5 agent rows6 kernel problemshard10 audit flags trackedDisplay only

The published KernelBench snapshot places Claude Fable 5 first at 23.5%. The third row is 2.5 points behind. The broader top-10 range is 10.1 points, so the table still separates the published systems.

5 models have been evaluated on KernelBench. The benchmark falls in the Coding category. KernelBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About KernelBench

Year

2026

Tasks

6 GPU-kernel optimization problems

Format

Mean peak fraction of hardware roofline over valid cells

Difficulty

Agentic GPU systems engineering

The mirrored board comes from the independent kernelbench.com Hard suite, not the Stanford KernelBench project. One long-running agent session tackles each problem. The visible score averages peak fraction of roofline over valid cells; failed and audit-flagged cells are excluded from that mean. We keep the benchmark display only because harness, hardware, and audit state are inseparable from the score.

Freshness and provenance

Version

KernelBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does KernelBench measure?

An agentic GPU-kernel benchmark that measures how much of the hardware roofline a model's correct, audit-clean kernels reach on six demanding CUDA and Triton problems.

Which model leads the published KernelBench snapshot?

Claude Fable 5 currently leads the published KernelBench snapshot with 23.5% mean peak fraction of roofline. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on KernelBench?

The September 23, 2026 snapshot snapshot contains 5 AI models.

Last updated: September 23, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.