Agent & Tool-Use Benchmarks
Atria Dawn Preview (Shanghai Artificial Intelligence Laboratory) leads BenchLM's verified agentic ranking at 92.5/100 across Terminal-Bench 2.0, OSWorld-Verified, and BrowseComp. The best open-weight agent model is Atria Dawn Preview at 92.5. This page catalogs all 26 agent benchmarks BenchLM tracks; the agentic category page carries the provisional and verified leaderboards.
Which AI models handle function calling, MCP tool use, browsing, and multi-step agent workflows best?Agentic work has its own BenchAlign v5.8 leaderboard. The overall ranking combines external indices with benchmark evidence rather than fixed category weights.
Only core agentic rows with an attached exact source record are shown. Source-unverified manual rows are excluded from the displayed agentic score and table cells.
- Best agentic model
- Atria Dawn PreviewVerified score 92.5 · Shanghai Artificial Intelligence Laboratory
- Best open-weight agent
- Atria Dawn PreviewVerified score 92.5 · Shanghai Artificial Intelligence Laboratory
- Benchmarks tracked
- 26 benchmarksTerminal, browsing, tool-use, and computer-use
Benchmark Categories
Core Weighted (3)
These 3 benchmarks determine agentic rankings
Tool Calling & MCP (6)
Function calling, MCP tool use, and structured workflows
Agent Frameworks (8)
OpenClaw-style and end-to-end agent evaluations
Computer & Browser Use (5)
Desktop GUI, mobile, and browser navigation tasks
Specialized (4)
Domain-specific agentic tasks across ML, research, and airline
Top 15 models by weighted agentic score
83 models match these filters.
Agentic score = weighted average of Terminal-Bench 2.0 (40%), OSWorld-Verified (35%), and BrowseComp (25%), normalized by available weights. This page intentionally stays on BenchLM's verified ranking lane and only includes exact-source rows. Display-only benchmarks (MCP Atlas, Toolathlon, etc.) are tracked but do not affect rankings.
Questions
What are LLM agent benchmarks?
Agent benchmarks test whether AI models can go beyond answering questions and actually complete multi-step tasks: browsing the web, writing and running code in a terminal, calling external APIs via function calling, and operating desktop or mobile interfaces. They measure real-world usefulness for autonomous workflows.
What is function calling and why does it matter?
Function calling (or tool use) lets an LLM invoke external tools, APIs, or databases as part of its response. This is critical for building AI agents that can search the web, query databases, send emails, or control other software. Benchmarks like BFCL v4 and Toolathlon specifically measure how reliably models select the right function and pass correct arguments.
What is MCP (Model Context Protocol)?
MCP is an open standard for connecting LLMs to external tools and data sources. MCP Atlas and MCP-Tasks benchmark how well models work with MCP-backed integrations. Strong MCP performance means a model integrates well into tool-rich agent architectures.
How does agentic work count in BenchLM scores?
Agentic work has its own BenchAlign v5.8 leaderboard. The overall ranking combines external indices with benchmark evidence rather than fixed category weights. The ability to use tools, browse, and complete multi-step tasks is the strongest differentiator between models in production use. A model that scores well on knowledge but cannot reliably call functions or navigate software has limited real-world utility for agent workflows.
Which models are best for building AI agents?
Currently, Atria Dawn Preview by Shanghai Artificial Intelligence Laboratory leads BenchLM's verified agentic rankings with a score of 92.5. The best open-weight agent model is Atria Dawn Preview (92.5). Check the leaderboard above for the full verified ranking.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.