Skip to main content
BenchLM

Agents Last Exam (ALE-Bench)

We show this table for reference; we do not rank on it.

A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.

Pass rate on ALE-Bench — June 2026 API snapshot

We mirror the published pass rate view for ALE-Bench. claude_code (thinking-max) / claude-opus-5-5 leads the public snapshot at 38.2%, followed by codex (reasoning-max) / gpt-6-astra (34.2%) and codex (reasoning-xhigh) / gpt-6-astra (32.2%). We do not use these results to rank models overall.

85 modelsAgenticCurrentDisplay onlyUpdated June 2026 API snapshot

Pass rate table (85 models)

Score
11
30.9%
24
28.3%
27
27.6%
29
27.0%
30
27.0%
31
27.0%
37
codex / gpt-5-5codexcodex
24.5%
42
ale_claw / gpt-5-5ale_clawale_claw
23.0%
45
22.4%
47
agnes_harness / Agnes-2.5-Pro-Betaagnes_harnessagnes_harness
21.7%
48
openclaw / gpt-5-5openclawopenclaw
21.7%
49
cursor_cli / gpt-5-5cursor_clicursor_cli
21.4%
52
openclaw / gpt-5-4openclawopenclaw
20.5%
54
20.4%
55
claude_code (max) / glm-5-2claude_codeclaude_code
20.4%
57
cursor_cli / composer-2-5cursor_clicursor_cli
20.4%
58
droid / gpt-5-5droiddroid
19.7%
59
claude_code / ark-0614cclaude_codeclaude_code
19.1%
60
18.4%
64
claude_code / claude-opus-4-8claude_codeclaude_code
16.4%
66
16.4%
67
15.8%
68
14.1%
70
claude_code / claude-opus-4-7claude_codeclaude_code
13.8%
71
13.5%
72
openclaw_cli / ark-0614copenclaw_cliopenclaw_cli
12.5%
73
12.4%
74
11.8%
76
ale_claw / gpt-5-4ale_clawale_claw
11.8%
77
openclaw / glm-5-1openclawopenclaw
11.5%
78
openclaw / kimi-k2-6openclawopenclaw
9.2%
79
openclaw / qwen3-6-plusopenclawopenclaw
8.6%
80
openclaw / mimo-v2-5openclawopenclaw
8.6%
81
grok_cli / grok-4-3grok_cligrok_cli
7.2%
82
codex / gpt-5-4codexcodex
7.2%
83
openclaw / minimax-m2-7openclawopenclaw
5.9%
84
grok_cli / grok-3grok_cligrok_cli
4.6%
85
openclaw / grok-4-3openclawopenclaw
4.3%

How ALE-Bench is shown here

BenchLM mirrors the Agents Last Exam full split from the public leaderboard API. The snapshot reports pass rate, partial average score, cost, token, and duration metadata across 152 professional workflow tasks for model plus agent-harness rows.

ALE-Bench is display only on BenchLM. Its rows combine a base model with an agent harness such as Codex, OpenClaw, Claude Code, Droid, Cursor CLI, or Gemini CLI, so BenchLM keeps the table separate from model-only rankings.

The Agent Showdown analysis adds domain and failure-mode context across 13 top-level domains. It also notes that Claude Code plus Fable 5 may include fallback to Opus 4.8 on refused tasks, so BenchLM preserves the official mixed-system row label instead of treating it as a pure base-model score.

Snapshot

85 harness rows152 ALE-V1 tasks13 domainsFull splitOfficial API snapshotDisplay only

The published ALE-Bench snapshot places claude_code (thinking-max) / claude-opus-5-5 first at 38.2%. The third row is 6.0 points behind. The broader top-10 range is 6.6 points, so many of the published results sit in a relatively narrow band.

85 models have been evaluated on ALE-Bench. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. ALE-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ALE-Bench

Year

2026

Tasks

152 ALE-V1 professional workflow tasks across 13 top-level domains

Format

Pass rate, partial-credit score, cost, token, and duration metadata

Difficulty

Real-world agentic workflows

BenchLM mirrors the public Agents Last Exam full leaderboard API as ALE-Bench and links the June 2026 Agent Showdown analysis for domain, cost, speed, and failure-mode context. Rows combine base models with agent harnesses such as Codex, OpenClaw, Claude Code, Droid, Cursor CLI, and Gemini CLI, so the table remains display-only. The source notes that Claude Code plus Fable 5 may include upstream fallback to Opus 4.8 on refused tasks.

Freshness and provenance

Version

ALE-Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does ALE-Bench measure?

A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.

Which model leads the published ALE-Bench snapshot?

claude_code (thinking-max) / claude-opus-5-5 currently leads the published ALE-Bench snapshot with 38.2% pass rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on ALE-Bench?

The June 2026 API snapshot snapshot contains 85 AI models.

Last updated: June 2026 API snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.