KernelBench Hard H100 (KernelBench)
We show this table for reference; we do not rank on it.
An agentic GPU-kernel benchmark that measures how much of the hardware roofline a model's correct, audit-clean kernels reach on six demanding CUDA and Triton problems.
Mean peak fraction of roofline on KernelBench — September 23, 2026 snapshot
We mirror the published mean peak fraction of roofline view for KernelBench. Claude Fable 5 leads the public snapshot at 23.5%, followed by Claude Opus 5 (21.8%) and Kimi K3 (20.9%). We do not use these results to rank models overall.
Claude Fable 5
Anthropic
claude · max
Claude Opus 5
Anthropic
or-opus · max
Kimi K3
Moonshot AI
kinetic-claude
5 modelsCodingCurrentDisplay onlyUpdated September 23, 2026 snapshot
Mean peak fraction of roofline table (5 models)
ScoreHow we show KernelBench
We mirror the KernelBench Hard hard board from kernelbench.com. The visible score is per-problem · rank by passes · bar = share of best on that problem across 6 GPU-kernel problems; failed and audit-flagged cells do not enter that mean.
This is an independent agent benchmark, not the Stanford KernelBench project leaderboard. Rows combine a model, an agent harness, long-running GPU sessions, and human audit decisions, so the table stays display only.
Snapshot
The published KernelBench snapshot places Claude Fable 5 first at 23.5%. The third row is 2.5 points behind. The broader top-10 range is 10.1 points, so the table still separates the published systems.
5 models have been evaluated on KernelBench. The benchmark falls in the Coding category. KernelBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About KernelBench
Year
2026
Tasks
6 GPU-kernel optimization problems
Format
Mean peak fraction of hardware roofline over valid cells
Difficulty
Agentic GPU systems engineering
The mirrored board comes from the independent kernelbench.com Hard suite, not the Stanford KernelBench project. One long-running agent session tackles each problem. The visible score averages peak fraction of roofline over valid cells; failed and audit-flagged cells are excluded from that mean. We keep the benchmark display only because harness, hardware, and audit state are inseparable from the score.
Freshness and provenance
Version
KernelBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does KernelBench measure?
An agentic GPU-kernel benchmark that measures how much of the hardware roofline a model's correct, audit-clean kernels reach on six demanding CUDA and Triton problems.
Which model leads the published KernelBench snapshot?
Claude Fable 5 currently leads the published KernelBench snapshot with 23.5% mean peak fraction of roofline. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on KernelBench?
The September 23, 2026 snapshot snapshot contains 5 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.