AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement (AI4AI-Bench)
Tests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation.
How to read this leaderboard
Each task maps its native metric onto a common scale where 0 means no useful learning, 0.1 matches the repository's original method, and 1 represents a perfect result. Scores above 0.1 improve on the shipped algorithm; scores below it made the method worse.
Operator receipt: 29 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5 at 0.288.
Honest limit: Each row combines model weights, reasoning effort, coding client, exploration budget, and ten different training environments. The normalized mean supports configuration comparison, not a controlled base-model rank.
How we show AI4AI-Bench
We mirror the benchmark-owned AI4AI-Bench configuration table from August 25, 2026 snapshot. The source reports 29 configurations across 10 AI research codebases, for 290 four-hour agent runs.
Each task maps its native held-out metric onto one scale: 0 means no useful learning, 0.1 matches the repository's original method, and 1 is perfect. We keep reasoning-effort and client configurations separate because the source result belongs to the full research-agent setup.
Snapshot
Mean normalized score on AI4AI-Bench — August 25, 2026 snapshot
We mirror the published mean normalized score view for AI4AI-Bench. Claude Opus 5 leads the public snapshot at 0.288, followed by Claude Opus 5 (0.278) and Claude Opus 5 (0.272). We do not use these results to rank models overall.
Claude Opus 5
Anthropic
Claude Code · medium reasoning
$166.31 exploration · 1.20M output tokens
Claude Opus 5
Anthropic
Claude Code · high reasoning
$181.09 exploration · 1.21M output tokens
Claude Opus 5
Anthropic
Claude Code · low reasoning
$181.27 exploration · 1.22M output tokens
Mean normalized score table (29 configurations)
ScoreThe published AI4AI-Bench snapshot places Claude Opus 5 first at 0.288. The third row is 0.016 score units behind. The broader top-10 range is 0.093 score units, so many of the published results sit in a relatively narrow band.
29 configurations have been evaluated on AI4AI-Bench. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. AI4AI-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About AI4AI-Bench
Year
2026
Tasks
10 AI training-algorithm design tasks
Format
Four-hour code rewrite followed by sealed training and held-out evaluation
Difficulty
End-to-end AI research and algorithm design
AI4AI-Bench freezes 10 research codebases spanning supervised fine-tuning, agentic reinforcement learning, distillation, reward modeling, preference optimization, diffusion reinforcement learning, machine unlearning, graph diffusion, model merging, and pruning. Agents receive four hours and one GPU to rewrite the method. A sealed pipeline then trains the submission from scratch for up to 12 hours and scores it on a held-out test. We mirror all 29 model, client, and reasoning-effort configurations as display-only evidence.
BenchLM freshness & provenance
Version
AI4AI-Bench v1.5
Refresh cadence
Quarterly
Staleness state
Current
Question availability
10 public tasks with score data and trajectories
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does AI4AI-Bench measure?
Tests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation.
Which model leads the published AI4AI-Bench snapshot?
Claude Opus 5 currently leads the published AI4AI-Bench snapshot with 0.288 mean normalized score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on AI4AI-Bench?
The August 25, 2026 snapshot contains 29 configurations across 6 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.