Agents Last Exam (ALE-Bench)
We show this table for reference; we do not rank on it.
A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.
Pass rate on ALE-Bench — June 2026 API snapshot
We mirror the published pass rate view for ALE-Bench. claude_code (thinking-max) / claude-opus-5-5 leads the public snapshot at 38.2%, followed by codex (reasoning-max) / gpt-6-astra (34.2%) and codex (reasoning-xhigh) / gpt-6-astra (32.2%). We do not use these results to rank models overall.
claude_code (thinking-max) / claude-opus-5-5
claude_code
claude_code
codex (reasoning-max) / gpt-6-astra
codex
codex
codex (reasoning-xhigh) / gpt-6-astra
codex
codex
85 modelsAgenticCurrentDisplay onlyUpdated June 2026 API snapshot
Pass rate table (85 models)
ScoreHow ALE-Bench is shown here
BenchLM mirrors the Agents Last Exam full split from the public leaderboard API. The snapshot reports pass rate, partial average score, cost, token, and duration metadata across 152 professional workflow tasks for model plus agent-harness rows.
ALE-Bench is display only on BenchLM. Its rows combine a base model with an agent harness such as Codex, OpenClaw, Claude Code, Droid, Cursor CLI, or Gemini CLI, so BenchLM keeps the table separate from model-only rankings.
The Agent Showdown analysis adds domain and failure-mode context across 13 top-level domains. It also notes that Claude Code plus Fable 5 may include fallback to Opus 4.8 on refused tasks, so BenchLM preserves the official mixed-system row label instead of treating it as a pure base-model score.
Snapshot
The published ALE-Bench snapshot places claude_code (thinking-max) / claude-opus-5-5 first at 38.2%. The third row is 6.0 points behind. The broader top-10 range is 6.6 points, so many of the published results sit in a relatively narrow band.
85 models have been evaluated on ALE-Bench. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. ALE-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About ALE-Bench
Year
2026
Tasks
152 ALE-V1 professional workflow tasks across 13 top-level domains
Format
Pass rate, partial-credit score, cost, token, and duration metadata
Difficulty
Real-world agentic workflows
BenchLM mirrors the public Agents Last Exam full leaderboard API as ALE-Bench and links the June 2026 Agent Showdown analysis for domain, cost, speed, and failure-mode context. Rows combine base models with agent harnesses such as Codex, OpenClaw, Claude Code, Droid, Cursor CLI, and Gemini CLI, so the table remains display-only. The source notes that Claude Code plus Fable 5 may include upstream fallback to Opus 4.8 on refused tasks.
Freshness and provenance
Version
ALE-Bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does ALE-Bench measure?
A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.
Which model leads the published ALE-Bench snapshot?
claude_code (thinking-max) / claude-opus-5-5 currently leads the published ALE-Bench snapshot with 38.2% pass rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on ALE-Bench?
The June 2026 API snapshot snapshot contains 85 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.