# WebArena-Verified Browser Agent Benchmark (WebArena-Verified)

> WebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task descriptions, reference answers, and evaluators, and replaces nondeterministic judging with deterministic checks where possible.

Canonical page: https://benchlm.ai/benchmarks/webarena-verified

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About WebArena-Verified

- Year: 2025
- Tasks: 812 verified tasks; separate 258-task Hard subset
- Format: Deterministic end-state task success
- Difficulty: Audited stateful browser work
- Paper: [WebArena-Verified: A Fully Audited Benchmark for Web Agents](https://openreview.net/forum?id=94tlGxmqkN)

The full release contains 812 verified tasks across six self-hosted web environments. A separate 258-task Hard subset supports lower-cost evaluation. The maintainers manually reviewed every task, reference answer, and evaluator, and removed LLM-as-a-judge and substring-matching checks in favor of deterministic scoring.

WebArena-Verified is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 69% |
| 2 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 66.8% |
| 3 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 64.8% |

## FAQ

### What changed in WebArena-Verified?

The maintainers manually reviewed and corrected every task, reference answer, and evaluator. They also replaced LLM judging and substring matching with deterministic, type-aware checks where possible. The full release has 812 tasks, with a separate 258-task Hard subset.

### Are WebArena-Verified scores directly comparable?

Only when the dataset subset, environment revision, agent scaffold, prompts, observation and action interfaces, tool or step budget, evaluator, and attempt policy match. A score belongs to that complete evaluation setup, not the base model alone.

### Does WebArena-Verified replace live workflow testing?

No. Its self-hosted environments support repeatable evaluation, but they do not reproduce every current website, login flow, anti-bot system, permission boundary, latency constraint, safety control, or operating cost.

## Compare Top Models on WebArena-Verified

- [Muse Spark 1.1 vs Qwen3.8 Max](/compare/muse-spark-1-1-vs-qwen3-8-max)
- [Qwen3.8 Max vs Qwen3.8-27B](/compare/qwen3-8-27b-vs-qwen3-8-max)
