# Artificial Analysis Coding Agent Index (AA Coding Agents)

> A display-only Artificial Analysis leaderboard for coding-agent systems, combining agent harnesses, host models, and execution settings across software-engineering benchmarks.

Canonical page: https://benchlm.ai/benchmarks/aacodingagents

- Category: [Coding](/coding)
- Last updated: Source captured 2026-09-23; evaluation dates not inferred

## About AA Coding Agents

- Year: 2026
- Tasks: Composite over DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA
- Format: Average pass@1 index
- Difficulty: Real-world coding-agent workflows
- Paper: [Artificial Analysis Coding Agent Benchmarks](https://artificialanalysis.ai/agents/coding-agents)

BenchLM mirrors the Artificial Analysis Coding Agent Index v1.1 page as a display-only external leaderboard. The source combines DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA component scores and publishes cost, token, and execution-time metadata. Rows are coding-agent systems rather than pure base-model results.

AA Coding Agents is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (19 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Code - Fable 5.1 (max) (with fallback)](https://artificialanalysis.ai/agents/coding-agents) | Anthropic | 62.2% |
| 2 | [Devin Fusion CLI - Claude Fable 5.1 XHigh + SWE-2 Medium](https://artificialanalysis.ai/agents/coding-agents) | Devin Fusion CLI | 61.7% |
| 3 | [Codex - GPT-6 Astra (max)](https://artificialanalysis.ai/agents/coding-agents) | OpenAI | 61.6% |
| 4 | [Claude Code - Opus 5 (max)](https://artificialanalysis.ai/agents/coding-agents) | Anthropic | 59.7% |
| 5 | [Devin Fusion CLI - GPT-6 Astra XHigh + SWE-2 Medium](https://artificialanalysis.ai/agents/coding-agents) | Devin Fusion CLI | 58.9% |
| 6 | [Codex - GPT-6 Sol (max)](https://artificialanalysis.ai/agents/coding-agents) | OpenAI | 56.7% |
| 7 | [Grok Build - Grok 4.7 (xhigh)](https://artificialanalysis.ai/agents/coding-agents) | xAI | 56.3% |
| 8 | [Codex - GPT-5.6 Sol (max) ({'reasoning_effort': 'max'})](https://artificialanalysis.ai/agents/coding-agents) | OpenAI | 54.6% |
| 9 | [Muse Code - Muse Spark 1.3 (max)](https://artificialanalysis.ai/agents/coding-agents) | Meta | 54.3% |
| 10 | [Opencode - GLM-5.3 ({'reasoning_effort': 'max'})](https://artificialanalysis.ai/agents/coding-agents) | Opencode | 53.6% |
| 11 | [Kimi Code CLI - Kimi K3](https://artificialanalysis.ai/agents/coding-agents) | Moonshot AI | 51.9% |
| 12 | [Muse Code - Muse Spark 1.3 (xhigh)](https://artificialanalysis.ai/agents/coding-agents) | Meta | 48.3% |
| 13 | [Grok Build - Grok 4.6 (xhigh)](https://artificialanalysis.ai/agents/coding-agents) | xAI | 47.0% |
| 14 | [Claude Code - Qwen3.8 Max](https://artificialanalysis.ai/agents/coding-agents) | Anthropic | 43.3% |
| 15 | [Codex - GPT-5.6 Luna (max)](https://artificialanalysis.ai/agents/coding-agents) | OpenAI | 43.2% |
| 16 | [Codex - DeepSeek V4 Pro 0813 (max)](https://artificialanalysis.ai/agents/coding-agents) | OpenAI | 43.1% |
| 17 | [Antigravity SDK - Gemini 3.8 Flash (high)](https://artificialanalysis.ai/agents/coding-agents) | Google | 41.9% |
| 18 | [Codex - GPT-6 Luna (max) ({'reasoning_effort': 'max'})](https://artificialanalysis.ai/agents/coding-agents) | OpenAI | 41.1% |
| 19 | [Codex - DeepSeek V4 Flash 0731 (max)](https://artificialanalysis.ai/agents/coding-agents) | OpenAI | 38.7% |

## FAQ

### What does AA Coding Agents measure?

A display-only Artificial Analysis leaderboard for coding-agent systems, combining agent harnesses, host models, and execution settings across software-engineering benchmarks.

### Which model leads the published AA Coding Agents snapshot?

Claude Code - Fable 5.1 (max) (with fallback) currently leads the published AA Coding Agents snapshot with a score of 62.2%.

### How many models are evaluated on AA Coding Agents?

The Source captured 2026-09-23; evaluation dates not inferred contains 19 AI models.

### Does AA Coding Agents affect BenchLM's overall score?

Not directly. AA Coding Agents is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
