Best LLMs for Agentic — September 2026 Leaderboard
Data refreshed:
Tool use, browser research, and computer-use workflows
As of September 2026, the top agentic model on the BenchLM leaderboard is Claude Fable 5.1 with a BenchAlign agentic score of 80.2.
Decision lens: the score determines position; Supported and Estimated labels describe the evidence behind that position without removing sparsely reported models.
- Data refreshed
- September 10, 2026
- Ranked
- 152 of 489 models
- Supported / Estimated
- 48 / 104
- Weighted evidence
- 4 of 30 benchmarks
30 tracked benchmarks
Terminal-Bench 2.0, BrowseComp, OSWorld-Verified, OSWorld 2.0, CyberGym, CWE-Bench, Cybench, ExploitGym, JobBench, BrowseComp-VL, OSWorld, AndroidWorld, WebVoyager, MCP Atlas, Toolathlon, Finance Agent v2, GDPval-AA, ZClawBench, Tau2-Telecom, DeepSearchQA, Tau2-Airline, PinchBench, OpenHands Index, SWE-Atlas Refactoring, SWE Refactor Bench, AI4AI-Bench, BFCL v4, MLE-Bench Lite, MM-ClawBench, Gert Labs
Scope: Terminal/tool use, Browser research, Computer use
Evidence set: Terminal-Bench 2.0, BrowseComp, OSWorld-Verified, OSWorld 2.0, CyberGym, CWE-Bench, Cybench, ExploitGym, JobBench, BrowseComp-VL, OSWorld, AndroidWorld, WebVoyager, MCP Atlas, Toolathlon, Finance Agent v2, GDPval-AA, ZClawBench, Tau2-Telecom, DeepSearchQA, Tau2-Airline, PinchBench, OpenHands Index, SWE-Atlas Refactoring, SWE Refactor Bench, AI4AI-Bench, BFCL v4, MLE-Bench Lite, MM-ClawBench, Gert Labs
Scope: Terminal/tool use, Browser research, Computer use
Best Agentic picks
BenchLM summaries for agentic plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.
Agentic AI Leaderboard
Primary score: BenchAlign agentic score. Higher values rank first. Use the Show metric control to change the value shown in each row.
Supported positions have diverse direct evidence. Estimated positions remain ranked but carry wider uncertainty.
Filters
1 | 80.2% | 80.16 | — | — | — | 41.7% | — | — | — | — | — | — | — | — | — | — | — | — | 1764 | — | — | — | — | — | — | — | — | — | — | — | — | — |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
2 | 78.1% | 78.07 | — | 90.8% | — | 70.6% | — | — | — | — | — | — | — | — | — | 85.8% | — | — | 1862 | — | — | 95.0% | — | — | — | — | — | — | — | — | — | — |
3 | 74.6% | 74.57 | 84.3% | — | 85% | — | — | — | — | — | — | — | — | — | — | — | — | — | 1747 | — | 98.5% | — | — | — | — | — | — | — | — | — | — | — |
4 | 71.9% | 71.91 | 88.3% | 91.2% | — | — | — | — | — | — | 52.9% | — | — | — | — | 84.2% | — | — | 1584 | — | — | 95.0% | — | — | — | — | — | — | — | — | — | — |
5 | 70.4% | 70.38 | — | 91.5% | — | 72.6% | — | — | — | 42.4% | — | — | — | — | — | — | — | — | 1580 | — | — | — | — | — | — | — | — | — | — | — | — | — |
6 | 70.0% | 70.05 | 91.9% | 92.2% | — | 62.6% | 84.5% | — | — | 33.7% | — | — | — | — | — | — | 58% | — | 1735 | — | 85.1% | — | — | — | — | — | — | — | — | — | — | — |
7 | 69.5% | 69.47 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 1643 | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 68.4% | 68.42 | — | — | — | — | 84.5% | — | — | 15.0% | — | — | — | — | — | — | — | — | 1769 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
9 | 67.2% | 67.16 | — | — | 86.1% | 19.4% | — | — | — | — | 53.4% | — | — | 85.3% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
10 | 66.3% | 66.31 | — | — | — | 59.0% | — | — | — | — | — | — | — | — | — | — | — | 61.4% | 1545 | — | — | — | — | — | — | — | — | — | — | — | — | — |
11 | 65.8% | 65.82 | 80.4% | 84.7% | 81.2% | — | — | — | — | — | — | — | — | — | — | — | — | — | 1603 | — | — | — | — | — | — | — | — | — | — | — | — | — |
12 | 64.0% | 64.03 | — | — | — | 47.9% | — | — | — | — | — | — | — | — | — | — | — | — | 1525 | — | — | — | — | — | — | — | — | — | — | — | — | — |
13 | 63.4% | 63.42 | — | — | 84.3% | — | — | — | — | — | 33.4% | — | — | 81.9% | — | — | — | — | 1463 | — | — | — | — | — | — | — | — | — | — | — | — | — |
14 | 62.5% | 62.47 | 74.6% | 84.3% | 83.4% | 20.6% | — | — | — | — | — | — | — | — | — | 82.2% | 59.9% | 53.9% | 1593 | — | 94.4% | 93.1% | — | — | — | — | — | — | — | — | — | 72.97% |
15 | 61.6% | 61.63 | 77.3% | — | 64.7% | — | — | — | — | — | 33.7% | — | — | — | — | — | — | — | — | — | 86% | — | — | — | — | — | — | — | — | — | — | 57.47% |
16 | 61.1% | 61.12 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 1631 | — | — | — | — | — | — | — | — | — | — | — | — | — |
17 | 61.0% | 61.04 | 82% | 84.4% | 78.7% | 13.0% | 81.8% | — | — | 13.4% | 42.7% | — | — | — | — | 75.3% | 55.6% | — | 1396 | — | 98% | — | — | — | — | — | — | — | — | — | — | 72.93% |
18 | 60.8% | 60.82 | — | 86.6% | — | — | — | — | — | — | — | — | — | — | — | 80% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
19 | 60.5% | 60.54 | 83.3% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 1430 | — | — | — | — | — | — | — | — | — | — | — | — | — |
20 | 60.2% | 60.2 | 87.4% | 87.5% | — | 50.2% | 81.8% | — | — | 23.2% | — | — | — | — | — | — | 53.1% | — | 1583 | — | 86.3% | — | — | — | — | — | — | — | — | — | — | — |
21 | 60.1% | 60.06 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 1773 | — | — | — | — | — | — | — | — | — | — | — | — | — |
22 | 59.7% | 59.74 | 65.4% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 95.9% | — | — | — | — | — | — | — | — | — | — | — |
23 | 59.4% | 59.39 | 80% | — | 80.8% | 14.2% | 59.0% | — | 92.9% | 0.8% | 54.7% | — | — | — | — | 88.1% | 75.6% | 57.2% | 1375 | — | — | 84.9% | — | — | — | — | — | — | — | — | — | — |
24 | 59.4% | 59.37 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 930 | — | 81.9% | — | — | — | — | — | — | — | — | — | — | 41.24% |
25 | 59.3% | 59.29 | — | 83.3% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 92.1% | — | — | — | — | — | — | — | — | — | — |
Top AI Models for Agentic — September 2026
As of September 2026, Claude Fable 5.1 leads the BenchAlign agentic leaderboard with a score of 80.2, followed by Claude Opus 5 (78.1) and Claude Fable 5 (74.6). BenchLM is currently showing 48 Supported and 104 Estimated models in this category.
Ranks #1 on the current agentic board with a Supported evidence label.
Ranks #2 on the current agentic board with a Supported evidence label.
Ranks #3 on the current agentic board with a Supported evidence label.
What changed
Claude Fable 5.1 ranks #1 at 80.2 with a Supported evidence label.
Claude Opus 5 ranks #2 at 78.1 with a Supported evidence label.
Claude Fable 5 ranks #3 at 74.6 with a Supported evidence label.
Top models by benchmark
Agentic software engineering and terminal task completion benchmark(30% of category score)
Score in Context
What these scores mean
BenchAlign places direct benchmarks and independent external signals on a common calibrated scale. The score is relative to the current evidence universe; it is not a raw percentage from any single test.
Known limitations
Estimated rows have less diverse direct evidence and wider uncertainty. They remain ranked so a newly released model is not treated as weak merely because fewer benchmark publishers have evaluated it.
How we weight
This lens combines category-relevant external evidence with admitted benchmark protocols. Evidence sources are calibrated for difficulty before aggregation, and no generated benchmark row contributes to the score.
Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.
The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.
Scroll horizontally to read the full evidence ledger.
| Benchmark | Weight | Status | Description |
|---|---|---|---|
| Terminal-Bench 2.0 | 30% | Weighted | Agentic software engineering and terminal task completion benchmark |
| BrowseComp | 25% | Weighted | Web research benchmark for browsing agents |
| OSWorld-Verified | 25% | Weighted | Computer-use benchmark for GUI task completion |
| OSWorld 2.0 | 10% | Weighted | A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks. |
| CyberGym | — | Display only | Cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance. |
| CWE-Bench | — | Display only | External benchmark for evaluating whether coding agents can produce correct patches for real-world software vulnerabilities. |
| Cybench | — | Display only | A cybersecurity benchmark of professional Capture the Flag tasks for measuring autonomous cyber agent capability and risk. |
| ExploitGym | — | Display only | A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits. |
| JobBench | — | Display only | An occupational agent benchmark for professional workflows that workers say they most want delegated to AI. |
| BrowseComp-VL | — | Display only | Vision-language browsing benchmark for multimodal web research and tool-use tasks. |
| OSWorld | — | Display only | Computer-use benchmark for GUI task completion across the broader OSWorld task suite. |
| AndroidWorld | — | Display only | Android GUI agent benchmark for task completion across mobile app workflows. |
| WebVoyager | — | Display only | Browser agent benchmark for completing multi-step workflows on live websites. |
| MCP Atlas | — | Display only | Tool-calling benchmark for Model Context Protocol integrations and multi-tool coordination |
| Toolathlon | — | Display only | General tool-calling benchmark for multi-step API and tool usage |
| Finance Agent v2 | — | Display only | Financial analysis and decision-making benchmark for agentic expert tasks. |
| GDPval-AA | — | Display only | Real-world agentic knowledge-work evaluation reported as an Elo score. |
| ZClawBench | — | Display only | Z.AI's OpenClaw workflow benchmark for broad agent tasks across research, office work, data analysis, devops, automation, and security. |
| Tau2-Telecom | — | Display only | Telecom-focused tool-use benchmark for structured API workflows |
| DeepSearchQA | — | Display only | Agentic browsing benchmark for list-style question answering with browser tools. |
| Tau2-Airline | — | Display only | Airline-domain tool-use benchmark for structured workflow execution and API correctness. |
| PinchBench | — | Display only | An OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows. |
| OpenHands Index | — | Display only | A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering. |
| SWE-Atlas Refactoring | — | Display only | A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks. |
| SWE Refactor Bench | — | Display only | Tests whether coding agents can complete long-horizon, whole-repository stack migrations while preserving the original program's behavior. |
| AI4AI-Bench | — | Display only | Tests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation. |
| BFCL v4 | — | Display only | Function-calling benchmark for tool selection, schema adherence, and argument correctness. |
| MLE-Bench Lite | — | Display only | A lightweight machine-learning competition benchmark that measures whether models can iteratively train, evaluate, and improve ML systems in low-resource settings. |
| MM-ClawBench | — | Display only | An OpenClaw-derived agent benchmark covering practical work and life tasks such as office document delivery, research, planning, and code maintenance. |
| Gert Labs | — | Display only | Composite game-environment leaderboard score across Gert Labs agentic coding, one-shot coding, and social decision-making modes. |
About Agentic Benchmarks
Agentic software engineering and terminal task completion benchmark
Common questions
What is an agentic LLM benchmark?
Agentic benchmarks evaluate whether AI models can complete multi-step workflows using tools, browsers, terminals, or software interfaces instead of only answering in chat.
Which benchmarks matter for AI agents?
Key agentic benchmarks include Terminal-Bench 2.0 for terminal tasks, BrowseComp for web research, and OSWorld-Verified for computer-use workflows.
Why do agentic benchmarks matter in 2026?
Agentic benchmarks matter because many modern products rely on models that can browse, plan, use tools, and complete end-to-end tasks rather than only generate text.
Agentic benchmark updates
Agentic is the fastest-moving category. Don't fall behind.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.