# Agents Last Exam (ALE-Bench)

> A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.

Canonical page: https://benchlm.ai/benchmarks/alebench

- Category: [Agentic](/agentic)
- Last updated: June 2026 API snapshot

## About ALE-Bench

- Year: 2026
- Tasks: 152 ALE-V1 professional workflow tasks across 13 top-level domains
- Format: Pass rate, partial-credit score, cost, token, and duration metadata
- Difficulty: Real-world agentic workflows
- Paper: [Agents Last Exam](https://agents-last-exam.org/leaderboard)

BenchLM mirrors the public Agents Last Exam full leaderboard API as ALE-Bench and links the June 2026 Agent Showdown analysis for domain, cost, speed, and failure-mode context. Rows combine base models with agent harnesses such as Codex, OpenClaw, Claude Code, Droid, Cursor CLI, and Gemini CLI, so the table remains display-only. The source notes that Claude Code plus Fable 5 may include upstream fallback to Opus 4.8 on refused tasks.

ALE-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (85 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [claude_code (thinking-max) / claude-opus-5-5](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 38.2% |
| 2 | [codex (reasoning-max) / gpt-6-astra](https://agents-last-exam.org/leaderboard) | codex | codex | 34.2% |
| 3 | [codex (reasoning-xhigh) / gpt-6-astra](https://agents-last-exam.org/leaderboard) | codex | codex | 32.2% |
| 4 | [codex (reasoning-medium) / gpt-6-astra](https://agents-last-exam.org/leaderboard) | codex | codex | 32.2% |
| 5 | [codex (reasoning-xhigh) / gpt-6-sol](https://agents-last-exam.org/leaderboard) | codex | codex | 32.2% |
| 6 | [claude_code (thinking-high) / claude-opus-5](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 32.2% |
| 7 | [codex (reasoning-xhigh) / muse-spark-1-3](https://agents-last-exam.org/leaderboard) | codex | codex | 32.2% |
| 8 | [codex (reasoning-max) / gpt-6-sol](https://agents-last-exam.org/leaderboard) | codex | codex | 31.6% |
| 9 | [codex (reasoning-high) / gpt-6-astra](https://agents-last-exam.org/leaderboard) | codex | codex | 31.6% |
| 10 | [codex (reasoning-high) / gpt-6-sol](https://agents-last-exam.org/leaderboard) | codex | codex | 31.6% |
| 11 | [claude_code (thinking-max) / claude-opus-5](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 30.9% |
| 12 | [codex (reasoning-xhigh) / GPT-5.6-Sol](https://agents-last-exam.org/leaderboard) | codex | codex | 30.6% |
| 13 | [codex (reasoning-high) / GPT-5.6-Sol](https://agents-last-exam.org/leaderboard) | codex | codex | 30.6% |
| 14 | [claude_code (thinking-xhigh) / claude-opus-5](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 30.3% |
| 15 | [codex (reasoning-xhigh) / GPT-5.6-Luna](https://agents-last-exam.org/leaderboard) | codex | codex | 30.3% |
| 16 | [codex (reasoning-max) / muse-spark-1-3](https://agents-last-exam.org/leaderboard) | codex | codex | 29.6% |
| 17 | [codex (reasoning-low) / gpt-6-astra](https://agents-last-exam.org/leaderboard) | codex | codex | 29.6% |
| 18 | [codex (reasoning-max) / GPT-5.6-Sol](https://agents-last-exam.org/leaderboard) | codex | codex | 29.6% |
| 19 | [codex (reasoning-medium) / GPT-5.6-Sol](https://agents-last-exam.org/leaderboard) | codex | codex | 29.5% |
| 20 | [codex (reasoning-medium) / gpt-6-sol](https://agents-last-exam.org/leaderboard) | codex | codex | 28.9% |
| 21 | [claude_code (thinking-medium) / claude-opus-5](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 28.9% |
| 22 | [codex (reasoning-max) / GPT-5.6-Luna](https://agents-last-exam.org/leaderboard) | codex | codex | 28.9% |
| 23 | [codex (reasoning-none) / gpt-6-astra](https://agents-last-exam.org/leaderboard) | codex | codex | 28.3% |
| 24 | [kimi_code (thinking-max) / kimi-k3](https://agents-last-exam.org/leaderboard) | kimi_code | kimi_code | 28.3% |
| 25 | [codex (reasoning-max) / GPT-5.6-Terra](https://agents-last-exam.org/leaderboard) | codex | codex | 28.0% |
| 26 | [codex (reasoning-low) / gpt-6-sol](https://agents-last-exam.org/leaderboard) | codex | codex | 27.6% |
| 27 | [claude_code (thinking-low) / claude-opus-5](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 27.6% |
| 28 | [codex (reasoning-xhigh) / GPT-5.6-Terra](https://agents-last-exam.org/leaderboard) | codex | codex | 27.6% |
| 29 | [claude_code (thinking-xhigh) / qwen-3-8-max](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 27.0% |
| 30 | [grok_build (reasoning-high) / grok-4-5](https://agents-last-exam.org/leaderboard) | grok_build | grok_build | 27.0% |
| 31 | [claude_code (thinking-max) / kimi-k3](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 27.0% |
| 32 | [claude_code (thinking-max) / claude-opus-4-8](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 27.0% |
| 33 | [codex (reasoning-xhigh) / gpt-5-5](https://agents-last-exam.org/leaderboard) | codex | codex | 26.6% |
| 34 | [codex (reasoning-high) / GPT-5.6-Terra](https://agents-last-exam.org/leaderboard) | codex | codex | 26.0% |
| 35 | [claude_code (reasoning-xhigh) / anthropic-claude-fable-5](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 25.7% |
| 36 | [codex (reasoning-max) / gpt-6-luna](https://agents-last-exam.org/leaderboard) | codex | codex | 25.0% |
| 37 | [codex / gpt-5-5](https://agents-last-exam.org/leaderboard) | codex | codex | 24.5% |
| 38 | [codex (reasoning-xhigh) / gpt-6-luna](https://agents-last-exam.org/leaderboard) | codex | codex | 24.3% |
| 39 | [codex (reasoning-high) / GPT-5.6-Luna](https://agents-last-exam.org/leaderboard) | codex | codex | 24.3% |
| 40 | [codex (reasoning-low) / GPT-5.6-Sol](https://agents-last-exam.org/leaderboard) | codex | codex | 23.9% |
| 41 | [codex (reasoning-none) / gpt-6-sol](https://agents-last-exam.org/leaderboard) | codex | codex | 23.7% |
| 42 | [ale_claw / gpt-5-5](https://agents-last-exam.org/leaderboard) | ale_claw | ale_claw | 23.0% |
| 43 | [codex (reasoning-medium) / GPT-5.6-Terra](https://agents-last-exam.org/leaderboard) | codex | codex | 22.7% |
| 44 | [claude_code (thinking-xhigh) / claude-opus-4-8](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 22.4% |
| 45 | [claude_code / anthropic-claude-fable-5](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 22.4% |
| 46 | [claude_code (thinking-high) / claude-opus-4-8](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 22.4% |
| 47 | [agnes_harness / Agnes-2.5-Pro-Beta](https://agents-last-exam.org/leaderboard) | agnes_harness | agnes_harness | 21.7% |
| 48 | [openclaw / gpt-5-5](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 21.7% |
| 49 | [cursor_cli / gpt-5-5](https://agents-last-exam.org/leaderboard) | cursor_cli | cursor_cli | 21.4% |
| 50 | [codex (reasoning-high) / gpt-6-luna](https://agents-last-exam.org/leaderboard) | codex | codex | 21.1% |
| 51 | [cursor_cli (thinking-high) / claude-opus-4-7](https://agents-last-exam.org/leaderboard) | cursor_cli | cursor_cli | 21.1% |
| 52 | [openclaw / gpt-5-4](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 20.5% |
| 53 | [codex (reasoning-medium) / gpt-6-luna](https://agents-last-exam.org/leaderboard) | codex | codex | 20.4% |
| 54 | [claude_code (reasoning-max) / qwen-3-8-27b](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 20.4% |
| 55 | [claude_code (max) / glm-5-2](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 20.4% |
| 56 | [codex (reasoning-low) / GPT-5.6-Terra](https://agents-last-exam.org/leaderboard) | codex | codex | 20.4% |
| 57 | [cursor_cli / composer-2-5](https://agents-last-exam.org/leaderboard) | cursor_cli | cursor_cli | 20.4% |
| 58 | [droid / gpt-5-5](https://agents-last-exam.org/leaderboard) | droid | droid | 19.7% |
| 59 | [claude_code / ark-0614c](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 19.1% |
| 60 | [ale_claw / claude-opus-4-7](https://agents-last-exam.org/leaderboard) | ale_claw | ale_claw | 18.4% |
| 61 | [codex (reasoning-medium) / gpt-5-5](https://agents-last-exam.org/leaderboard) | codex | codex | 18.4% |
| 62 | [codex (reasoning-medium) / GPT-5.6-Luna](https://agents-last-exam.org/leaderboard) | codex | codex | 17.1% |
| 63 | [codex (reasoning-low) / gpt-5-5](https://agents-last-exam.org/leaderboard) | codex | codex | 17.1% |
| 64 | [claude_code / claude-opus-4-8](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 16.4% |
| 65 | [codex (reasoning-low) / gpt-6-luna](https://agents-last-exam.org/leaderboard) | codex | codex | 16.4% |
| 66 | [gemini_cli / gemini-3-1-pro-preview](https://agents-last-exam.org/leaderboard) | gemini_cli | gemini_cli | 16.4% |
| 67 | [openclaw / claude-opus-4-7](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 15.8% |
| 68 | [openclaw / gemini-3-1-pro-preview](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 14.1% |
| 69 | [codex (reasoning-none) / gpt-6-luna](https://agents-last-exam.org/leaderboard) | codex | codex | 13.8% |
| 70 | [claude_code / claude-opus-4-7](https://agents-last-exam.org/leaderboard) | claude_code | claude_code | 13.8% |
| 71 | [droid / claude-opus-4-7](https://agents-last-exam.org/leaderboard) | droid | droid | 13.5% |
| 72 | [openclaw_cli / ark-0614c](https://agents-last-exam.org/leaderboard) | openclaw_cli | openclaw_cli | 12.5% |
| 73 | [openclaw / deepseek-v4-pro](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 12.4% |
| 74 | [openclaw / qwen-qwen3-7-max](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 11.8% |
| 75 | [codex (reasoning-low) / GPT-5.6-Luna](https://agents-last-exam.org/leaderboard) | codex | codex | 11.8% |
| 76 | [ale_claw / gpt-5-4](https://agents-last-exam.org/leaderboard) | ale_claw | ale_claw | 11.8% |
| 77 | [openclaw / glm-5-1](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 11.5% |
| 78 | [openclaw / kimi-k2-6](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 9.2% |
| 79 | [openclaw / qwen3-6-plus](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 8.6% |
| 80 | [openclaw / mimo-v2-5](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 8.6% |
| 81 | [grok_cli / grok-4-3](https://agents-last-exam.org/leaderboard) | grok_cli | grok_cli | 7.2% |
| 82 | [codex / gpt-5-4](https://agents-last-exam.org/leaderboard) | codex | codex | 7.2% |
| 83 | [openclaw / minimax-m2-7](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 5.9% |
| 84 | [grok_cli / grok-3](https://agents-last-exam.org/leaderboard) | grok_cli | grok_cli | 4.6% |
| 85 | [openclaw / grok-4-3](https://agents-last-exam.org/leaderboard) | openclaw | openclaw | 4.3% |

## FAQ

### What does ALE-Bench measure?

A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.

### Which model leads the published ALE-Bench snapshot?

claude_code (thinking-max) / claude-opus-5-5 currently leads the published ALE-Bench snapshot with a score of 38.2%.

### How many models are evaluated on ALE-Bench?

The June 2026 API snapshot contains 85 AI models.

### Does ALE-Bench affect BenchLM's overall score?

Not directly. ALE-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
