# LLM agent and tool-use benchmarks

Compare models across agentic evaluations covering tool use, software tasks, browsing, computer use, and multi-step execution.

Last updated: October 1, 2026

Canonical page: https://benchlm.ai/llm-agent-benchmarks
