Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement (AI4AI-Bench)

Tests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation.

How to read this leaderboard

Each task maps its native metric onto a common scale where 0 means no useful learning, 0.1 matches the repository's original method, and 1 represents a perfect result. Scores above 0.1 improve on the shipped algorithm; scores below it made the method worse.

Operator receipt: 29 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5 at 0.288.

Honest limit: Each row combines model weights, reasoning effort, coding client, exploration budget, and ten different training environments. The normalized mean supports configuration comparison, not a controlled base-model rank.

How we show AI4AI-Bench

We mirror the benchmark-owned AI4AI-Bench configuration table from August 25, 2026 snapshot. The source reports 29 configurations across 10 AI research codebases, for 290 four-hour agent runs.

Each task maps its native held-out metric onto one scale: 0 means no useful learning, 0.1 matches the repository's original method, and 1 is perfect. We keep reasoning-effort and client configurations separate because the source result belongs to the full research-agent setup.

Snapshot

6 models29 configurations10 research codebases290 agent runs124 below baselineDisplay only

Mean normalized score on AI4AI-Bench — August 25, 2026 snapshot

We mirror the published mean normalized score view for AI4AI-Bench. Claude Opus 5 leads the public snapshot at 0.288, followed by Claude Opus 5 (0.278) and Claude Opus 5 (0.272). We do not use these results to rank models overall.

29 configurationsAgenticCurrentDisplay onlyUpdated August 25, 2026 snapshot

Mean normalized score table (29 configurations)

Score
1
Claude Opus 5Anthropic · ClosedClaude Code · medium reasoning$166.31 exploration · 1.20M output tokens
0.288
2
Claude Opus 5Anthropic · ClosedClaude Code · high reasoning$181.09 exploration · 1.21M output tokens
0.278
3
Claude Opus 5Anthropic · ClosedClaude Code · low reasoning$181.27 exploration · 1.22M output tokens
0.272
4
GPT-5.6 SolOpenAI · ClosedCodex · max reasoning$626.19 exploration · 1.25M output tokens
0.245
5
GPT-5.6 SolOpenAI · ClosedCodex · high reasoning$419.93 exploration · 0.62M output tokens
0.235
6
GPT-5.6 TerraOpenAI · ClosedCodex · max reasoning$345.99 exploration · 1.57M output tokens
0.228
7
Claude Opus 5Anthropic · ClosedClaude Code · max reasoning$195.47 exploration · 1.34M output tokens
0.225
8
Claude Sonnet 5Anthropic · ClosedClaude Code · high reasoning$95.53 exploration · 1.03M output tokens
0.216
9
GPT-5.6 SolOpenAI · ClosedCodex · low reasoning$336.86 exploration · 0.44M output tokens
0.198
10
GPT-5.6 SolOpenAI · ClosedCodex · xhigh reasoning$521.20 exploration · 0.83M output tokens
0.195
11
Claude Opus 5Anthropic · ClosedClaude Code · xhigh reasoning$184.67 exploration · 1.34M output tokens
0.190
12
Kimi K3Moonshot AI · ClosedClaude Code · max reasoning$29.93 exploration · 0.61M output tokens
0.174
13
Claude Sonnet 5Anthropic · ClosedClaude Code · max reasoning$108.62 exploration · 1.10M output tokens
0.164
14
GPT-5.6 LunaOpenAI · ClosedCodex · xhigh reasoning$109.89 exploration · 0.71M output tokens
0.147
15
GPT-5.6 SolOpenAI · ClosedCodex · medium reasoning$448.69 exploration · 0.61M output tokens
0.144
16
GPT-5.6 TerraOpenAI · ClosedCodex · xhigh reasoning$228.57 exploration · 0.91M output tokens
0.144
17
GPT-5.6 LunaOpenAI · ClosedCodex · max reasoning$107.87 exploration · 0.77M output tokens
0.139
18
Claude Sonnet 5Anthropic · ClosedClaude Code · xhigh reasoning$100.78 exploration · 1.02M output tokens
0.137
19
GPT-5.6 TerraOpenAI · ClosedCodex · medium reasoning$43.34 exploration · 0.24M output tokens
0.135
20
GPT-5.6 TerraOpenAI · ClosedCodex · high reasoning$213.87 exploration · 0.73M output tokens
0.129
21
GPT-5.6 SolOpenAI · ClosedCodex · none reasoning$339.34 exploration · 0.35M output tokens
0.126
22
GPT-5.6 LunaOpenAI · ClosedCodex · medium reasoning$30.35 exploration · 0.26M output tokens
0.124
23
GPT-5.6 LunaOpenAI · ClosedCodex · high reasoning$65.93 exploration · 0.50M output tokens
0.119
24
Claude Sonnet 5Anthropic · ClosedClaude Code · medium reasoning$97.75 exploration · 0.91M output tokens
0.109
25
GPT-5.6 TerraOpenAI · ClosedCodex · low reasoning$34.56 exploration · 0.18M output tokens
0.108
26
Claude Sonnet 5Anthropic · ClosedClaude Code · low reasoning$93.07 exploration · 0.91M output tokens
0.099
27
GPT-5.6 LunaOpenAI · ClosedCodex · none reasoning$16.91 exploration · 0.12M output tokens
0.091
28
GPT-5.6 LunaOpenAI · ClosedCodex · low reasoning$6.51 exploration · 0.08M output tokens
0.079
29
GPT-5.6 TerraOpenAI · ClosedCodex · none reasoning$3.53 exploration · 0.02M output tokens
0.065

The published AI4AI-Bench snapshot places Claude Opus 5 first at 0.288. The third row is 0.016 score units behind. The broader top-10 range is 0.093 score units, so many of the published results sit in a relatively narrow band.

29 configurations have been evaluated on AI4AI-Bench. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. AI4AI-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AI4AI-Bench

Year

2026

Tasks

10 AI training-algorithm design tasks

Format

Four-hour code rewrite followed by sealed training and held-out evaluation

Difficulty

End-to-end AI research and algorithm design

AI4AI-Bench freezes 10 research codebases spanning supervised fine-tuning, agentic reinforcement learning, distillation, reward modeling, preference optimization, diffusion reinforcement learning, machine unlearning, graph diffusion, model merging, and pruning. Agents receive four hours and one GPU to rewrite the method. A sealed pipeline then trains the submission from scratch for up to 12 hours and scores it on a held-out test. We mirror all 29 model, client, and reasoning-effort configurations as display-only evidence.

BenchLM freshness & provenance

Version

AI4AI-Bench v1.5

Refresh cadence

Quarterly

Staleness state

Current

Question availability

10 public tasks with score data and trajectories

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AI4AI-Bench measure?

Tests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation.

Which model leads the published AI4AI-Bench snapshot?

Claude Opus 5 currently leads the published AI4AI-Bench snapshot with 0.288 mean normalized score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on AI4AI-Bench?

The August 25, 2026 snapshot contains 29 configurations across 6 AI models.

Last updated: August 25, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.