Skip to main content

Benchmark profile

Agents Last Exam (ALE-Bench)

A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.

How BenchLM shows ALE-Bench

BenchLM mirrors the Agents Last Exam full split from the public leaderboard API. The snapshot reports pass rate, partial average score, cost, token, and duration metadata across 152 professional workflow tasks for model plus agent-harness rows.

ALE-Bench is display only on BenchLM. Its rows combine a base model with an agent harness such as Codex, OpenClaw, Claude Code, Droid, Cursor CLI, or Gemini CLI, so BenchLM keeps the table separate from model-only rankings.

The Agent Showdown analysis adds domain and failure-mode context across 13 top-level domains. It also notes that Claude Code plus Fable 5 may include fallback to Opus 4.8 on refused tasks, so BenchLM preserves the official mixed-system row label instead of treating it as a pure base-model score.

53 harness rows152 ALE-V1 tasks13 domainsFull splitOfficial API snapshotDisplay only

Pass rate on ALE-Bench — June 2026 API snapshot

BenchLM mirrors the published pass rate view for ALE-Bench. codex (reasoning-xhigh) / GPT-5.6-Sol leads the public snapshot at 30.6% , followed by codex (reasoning-high) / GPT-5.6-Sol (30.6%) and codex (reasoning-max) / GPT-5.6-Sol (29.6%). BenchLM does not use these results to rank models overall.

53 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated June 2026 API snapshot

Pass rate table (53 models)

Score
13
24.2%
16
23.0%
21
21.1%
22
20.7%
23
20.5%
25
20.4%
27
20.4%
28
19.1%
29
19.1%
40
12.5%
44
11.8%
45
11.5%
46
9.2%
48
8.6%
49
7.2%
50
6.6%
52
4.6%
53
4.3%

The published ALE-Bench snapshot places codex (reasoning-xhigh) / GPT-5.6-Sol first at 30.6%. The third row is 1.0 points behind. The broader top-10 range is 4.0 points, so many of the published results sit in a relatively narrow band.

53 models have been evaluated on ALE-Bench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. ALE-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ALE-Bench

Year

2026

Tasks

152 ALE-V1 professional workflow tasks across 13 top-level domains

Format

Pass rate, partial-credit score, cost, token, and duration metadata

Difficulty

Real-world agentic workflows

BenchLM mirrors the public Agents Last Exam full leaderboard API as ALE-Bench and links the June 2026 Agent Showdown analysis for domain, cost, speed, and failure-mode context. Rows combine base models with agent harnesses such as Codex, OpenClaw, Claude Code, Droid, Cursor CLI, and Gemini CLI, so the table remains display-only. The source notes that Claude Code plus Fable 5 may include upstream fallback to Opus 4.8 on refused tasks.

BenchLM freshness & provenance

Version

ALE-Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ALE-Bench measure?

A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.

Which model leads the published ALE-Bench snapshot?

codex (reasoning-xhigh) / GPT-5.6-Sol currently leads the published ALE-Bench snapshot with 30.6% pass rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on ALE-Bench?

53 AI models are included in BenchLM's mirrored ALE-Bench snapshot, based on the public leaderboard captured on June 2026 API snapshot.

Last updated: June 2026 API snapshot · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.