Skip to main content
BenchLM

Agent & Tool-Use Benchmarks

Atria Dawn Preview (Shanghai Artificial Intelligence Laboratory) leads BenchLM's verified agentic ranking at 92.5/100 across Terminal-Bench 2.0, OSWorld-Verified, and BrowseComp. The best open-weight agent model is Atria Dawn Preview at 92.5. This page catalogs all 26 agent benchmarks BenchLM tracks; the agentic category page carries the provisional and verified leaderboards.

Benchmark data verified

Which AI models handle function calling, MCP tool use, browsing, and multi-step agent workflows best?Agentic work has its own BenchAlign v5.8 leaderboard. The overall ranking combines external indices with benchmark evidence rather than fixed category weights.

Only core agentic rows with an attached exact source record are shown. Source-unverified manual rows are excluded from the displayed agentic score and table cells.

Best agentic model
Atria Dawn Preview
Verified score 92.5 · Shanghai Artificial Intelligence Laboratory
Best open-weight agent
Atria Dawn Preview
Verified score 92.5 · Shanghai Artificial Intelligence Laboratory
Benchmarks tracked
26 benchmarks
Terminal, browsing, tool-use, and computer-use

Benchmark Categories

Core Weighted (3)

These 3 benchmarks determine agentic rankings

Tool Calling & MCP (6)

Function calling, MCP tool use, and structured workflows

Computer & Browser Use (5)

Desktop GUI, mobile, and browser navigation tasks

Specialized (4)

Domain-specific agentic tasks across ML, research, and airline

Top 15 models by weighted agentic score

OpenAIAnthropicGoogleMetaDeepSeekMistralxAIAlibaba

83 models match these filters.

CSVJSON
RankModelVerified score
1
Atria Dawn PreviewShanghai Artificial Intelligence Laboratory · open weightTB —BC 92.5OS —
92.5agentic / 100
2
GPT-5.6 SolOpenAI · proprietaryTB —BC 92.2OS —
92.2agentic / 100
3
GPT-6 AstraOpenAI · proprietaryTB —BC 91.5OS —
91.5agentic / 100
4
Kimi K3Moonshot AI · proprietaryTB —BC 91.2OS —
91.2agentic / 100
5
Claude Opus 5Anthropic · proprietaryTB —BC 90.8OS —
90.8agentic / 100
6
GPT-5.5 ProOpenAI · proprietaryTB —BC 90.1OS —
90.1agentic / 100
7
GPT-5.4 ProOpenAI · proprietaryTB —BC 89.3OS —
89.3agentic / 100
8
Step 5 PreviewStepFun · proprietaryTB —BC 88.7OS —
88.7agentic / 100
9
GPT-5.6 TerraOpenAI · proprietaryTB —BC 87.5OS —
87.5agentic / 100
10
Ornith-1.5-397BOrnith AI · open weightTB —BC 86.6OS —
86.6agentic / 100
12
Qwen3.8 MaxAlibaba · open weightTB —BC —OS 86.1
86.1agentic / 100
13
Claude Fable 5Anthropic · proprietaryTB —BC —OS 85
85agentic / 100
14
Qwen3.8-27BAlibaba · open weightTB —BC —OS 84.3
84.3agentic / 100
15
Claude Opus 4.8Anthropic · proprietaryTB —BC 84.3OS 83.4
83.9agentic / 100
16
GPT-5.6 LunaOpenAI · proprietaryTB —BC 83.3OS —
83.3agentic / 100
18
Claude Sonnet 5Anthropic · proprietaryTB —BC 84.7OS 81.2
83agentic / 100
20
Holo3-35B-A3BH Company · open weightTB —BC —OS 82.56
82.6agentic / 100
21
MiMo-V2.6-ProXiaomi · open weightTB —BC —OS 82
82agentic / 100
22
GPT-5.5OpenAI · proprietaryTB 82BC 84.4OS 78.7
81.7agentic / 100
23
Muse Spark 1.1Meta · proprietaryTB —BC —OS 80.8
80.8agentic / 100
25
Holo3-122B-A10BH Company · proprietaryTB —BC —OS 78.85
78.9agentic / 100
27
GPT-5.4OpenAI · proprietaryTB 75.1BC 82.7OS 75
77.4agentic / 100
28
Inkling-SmallThinking Machines Lab · open weightTB —BC 77.4OS —
77.4agentic / 100
29
InklingThinking Machines Lab · open weightTB —BC 77.1OS —
77.1agentic / 100
30
UI-Mate-27BTencent · open weightTB —BC —OS 77
77agentic / 100
31
MiniMax M3MiniMax · open weightTB —BC 83.52OS 70.06
76.8agentic / 100
32
Step 3.7 FlashStepFun · open weightTB —BC 75.82OS —
75.8agentic / 100
33
Agents-A1InternScience · open weightTB —BC 75.51OS —
75.5agentic / 100
37
Kimi K2.6Moonshot AI · open weightTB 66.7BC 83.2OS 73.1
73.9agentic / 100
38
Claude Opus 4.6Anthropic · proprietaryTB 65.4BC 83.7OS 72.7
73.4agentic / 100
39
Ling 3.0 FlashInclusionAI · open weightTB —BC 72.2OS —
72.2agentic / 100
40
Qwen3.7 PlusAlibaba · proprietaryTB 70.3BC —OS 73.3
71.7agentic / 100
41
GPT-5.3 CodexOpenAI · proprietaryTB 77.3BC —OS 64.7
71.6agentic / 100
42
Qwen3.7 MaxAlibaba · proprietaryTB 69.7BC —OS —
69.7agentic / 100
43
Composer 2.5Cursor · proprietaryTB 69.3BC —OS —
69.3agentic / 100
44
MiMo-V2.5-ProXiaomi · proprietaryTB 68.4BC —OS —
68.4agentic / 100
46
Agents-A1-4BInternScience · open weightTB —BC 66.8OS —
66.8agentic / 100
47
UI-Mate-9BTencent · open weightTB —BC —OS 66.2
66.2agentic / 100
49
MiMo-V2.5Xiaomi · proprietaryTB 65.8BC —OS —
65.8agentic / 100
50
GLM-5.1Z.AI · open weightTB 63.5BC 68OS —
65.5agentic / 100
51
GPT-5.4 miniOpenAI · proprietaryTB 60BC —OS 72.1
65.5agentic / 100
55
Claude Opus 4.5Anthropic · proprietaryTB 59.3BC —OS 66.3
62.5agentic / 100
56
Composer 2Cursor · proprietaryTB 61.7BC —OS —
61.7agentic / 100
57
Qwen3.6 PlusAlibaba · proprietaryTB 61.6BC —OS —
61.6agentic / 100
58
Qwen3.6-27BAlibaba · open weightTB 59.3BC —OS —
59.3agentic / 100
59
Muse SparkMeta · proprietaryTB 59BC —OS —
59agentic / 100
60
MiniMax M2.7MiniMax · open weightTB 57BC —OS —
57agentic / 100
61
Qwen3.5 397BAlibaba · open weightTB 52.5BC 62OS —
56.8agentic / 100
62
GPT-5.2OpenAI · proprietaryTB —BC 65.8OS 47.3
56.6agentic / 100
64
Ornith-1.5-9BOrnith AI · open weightTB —BC 56.4OS —
56.4agentic / 100
65
GLM-5Z.AI · open weightTB 56.2BC —OS —
56.2agentic / 100
66
Kimi K2.5Moonshot AI · open weightTB 50.8BC 60.6OS —
55.3agentic / 100
69
Hy3 PreviewTencent · open weightTB 54.4BC —OS —
54.4agentic / 100
70
Qwen3.5-27BAlibaba · open weightTB 41.6BC 61OS 56.2
52.2agentic / 100
71
Qwen3.6-35B-A3BAlibaba · open weightTB 51.5BC —OS —
51.5agentic / 100
72
Qwen3.5-35B-A3BAlibaba · open weightTB 40.5BC 61OS 54.5
51.3agentic / 100
73
Solar Pro 4Upstage · proprietaryTB —BC 49.2OS —
49.2agentic / 100
74
Grok 4.20xAI · proprietaryTB 47.1BC —OS —
47.1agentic / 100
75
GLM-4.7Z.AI · open weightTB 41BC 52OS —
46agentic / 100
76
MAI-Thinking-1Microsoft · proprietaryTB 46BC —OS —
46agentic / 100
77
Laguna M.1Poolside · proprietaryTB 45.8BC —OS —
45.8agentic / 100
79
GPT-5.4 nanoOpenAI · proprietaryTB 46.3BC —OS 39
43agentic / 100
81
Laguna XS 2.1Poolside · open weightTB 37.5BC —OS —
37.5agentic / 100
83
Laguna XS.2Poolside · open weightTB 35.7BC —OS —
35.7agentic / 100

Agentic score = weighted average of Terminal-Bench 2.0 (40%), OSWorld-Verified (35%), and BrowseComp (25%), normalized by available weights. This page intentionally stays on BenchLM's verified ranking lane and only includes exact-source rows. Display-only benchmarks (MCP Atlas, Toolathlon, etc.) are tracked but do not affect rankings.

Questions

What are LLM agent benchmarks?

Agent benchmarks test whether AI models can go beyond answering questions and actually complete multi-step tasks: browsing the web, writing and running code in a terminal, calling external APIs via function calling, and operating desktop or mobile interfaces. They measure real-world usefulness for autonomous workflows.

What is function calling and why does it matter?

Function calling (or tool use) lets an LLM invoke external tools, APIs, or databases as part of its response. This is critical for building AI agents that can search the web, query databases, send emails, or control other software. Benchmarks like BFCL v4 and Toolathlon specifically measure how reliably models select the right function and pass correct arguments.

What is MCP (Model Context Protocol)?

MCP is an open standard for connecting LLMs to external tools and data sources. MCP Atlas and MCP-Tasks benchmark how well models work with MCP-backed integrations. Strong MCP performance means a model integrates well into tool-rich agent architectures.

How does agentic work count in BenchLM scores?

Agentic work has its own BenchAlign v5.8 leaderboard. The overall ranking combines external indices with benchmark evidence rather than fixed category weights. The ability to use tools, browse, and complete multi-step tasks is the strongest differentiator between models in production use. A model that scores well on knowledge but cannot reliably call functions or navigate software has limited real-world utility for agent workflows.

Which models are best for building AI agents?

Currently, Atria Dawn Preview by Shanghai Artificial Intelligence Laboratory leads BenchLM's verified agentic rankings with a score of 92.5. The best open-weight agent model is Atria Dawn Preview (92.5). Check the leaderboard above for the full verified ranking.

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.