# WebArena Web Agent Benchmark (WebArena)

> WebArena tests whether a browser-agent system can complete 812 long-horizon tasks inside self-hosted replicas of functional websites. It checks the requested end state, so a result reflects the model, agent scaffold, browser interface, action budget, and evaluator together—not the base model alone.

Canonical page: https://benchlm.ai/benchmarks/webarena

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About WebArena

- Year: 2024
- Tasks: 812 long-horizon browser tasks
- Format: End-state task success
- Difficulty: Stateful multi-site browser work
- Paper: [WebArena: A Realistic Web Environment for Building Autonomous Agents](https://arxiv.org/abs/2307.13854)

Tasks start from high-level natural-language goals across shopping, forums, collaborative software development, content management, maps, and knowledge sites. Evaluation uses answer matching or programmatic checks of site state, allowing more than one valid action path. This route owns original WebArena methodology. The separately audited WebArena-Verified release has its own result route.

WebArena is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### What does WebArena measure?

WebArena measures whether a browser-agent system can complete 812 long-horizon tasks in self-hosted website replicas. Tasks cover shopping, forums, software collaboration, content management, maps, and knowledge work. Success is checked against the requested answer or resulting site state, so multiple valid action paths can pass.

### Are WebArena scores directly comparable?

No. Match original WebArena or WebArena-Verified, full suite or subset, environment revision, agent scaffold, prompts, observation and action interfaces, tool or step budget, evaluator, and attempt policy. A model tested with different browser tools or repeated attempts is not part of the same controlled comparison.

### Can WebArena choose the best browser agent?

Not by itself. WebArena is a useful signal for stateful browser work, but its fixed self-hosted sites do not reproduce today's changing web, production login flows, anti-bot systems, real-world identity and permission boundaries, latency, cost, or safety controls. Use matched WebArena results alongside live workflow trials and broader agentic evidence.
