Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start the free Radar Brief

WebArena-Verified Browser Agent Benchmark (WebArena-Verified)

WebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task descriptions, reference answers, and evaluators, and replaces nondeterministic judging with deterministic checks where possible.

Data verified 25 confirmed releases in the last 30 daysStart the free Radar Brief

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Treat each score as a result for the complete model-agent system, not the base model alone. Compare rows only when the full suite or Hard subset, environment revision, scaffold, prompts, DOM or screenshot interface, action and tool budget, evaluator, and pass@1 or repeat policy match.

Operator receipt: 3 sourced rows are currently displayable on this page; the leading published row is Muse Spark 1.1 at 69%.

Honest limit: The live ledger currently contains one exact provider-reported row, so it cannot establish a broad market leader. That run used the full 812-task suite, OpenClaw, hybrid DOM and screenshot interaction, a 150-tool-call cap, and mean pass@1. The fixed environments still omit much of today's open-web drift, production authentication, anti-bot behavior, permission boundaries, latency, safety controls, and cost.

Benchmark score on WebArena-Verified — August 29, 2026

We mirror the published score view for WebArena-Verified. Muse Spark 1.1 leads the public snapshot at 69%, followed by Qwen3.8 Max (66.8%) and Qwen3.8-27B (64.8%). We do not use these results to rank models overall.

3 modelsAgenticCurrentDisplay onlyUpdated August 29, 2026

Benchmark score table (3 models)

Score
1
Muse Spark 1.1Meta · Closed
69%
2
Qwen3.8 MaxAlibaba · Open weight
66.8%
3
Qwen3.8-27BAlibaba · Open weight
64.8%

The published WebArena-Verified snapshot places Muse Spark 1.1 first at 69%. The third row is 4.2 points behind. The broader top-10 range is 4.2 points, so many of the published results sit in a relatively narrow band.

3 models have been evaluated on WebArena-Verified. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. WebArena-Verified is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About WebArena-Verified

Year

2025

Tasks

812 verified tasks; separate 258-task Hard subset

Format

Deterministic end-state task success

Difficulty

Audited stateful browser work

The full release contains 812 verified tasks across six self-hosted web environments. A separate 258-task Hard subset supports lower-cost evaluation. The maintainers manually reviewed every task, reference answer, and evaluator, and removed LLM-as-a-judge and substring-matching checks in favor of deterministic scoring.

BenchLM freshness & provenance

Version

WebArena-Verified 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What changed in WebArena-Verified?

The maintainers manually reviewed and corrected every task, reference answer, and evaluator. They also replaced LLM judging and substring matching with deterministic, type-aware checks where possible. The full release has 812 tasks, with a separate 258-task Hard subset.

Are WebArena-Verified scores directly comparable?

Only when the dataset subset, environment revision, agent scaffold, prompts, observation and action interfaces, tool or step budget, evaluator, and attempt policy match. A score belongs to that complete evaluation setup, not the base model alone.

Does WebArena-Verified replace live workflow testing?

No. Its self-hosted environments support repeatable evaluation, but they do not reproduce every current website, login flow, anti-bot system, permission boundary, latency constraint, safety control, or operating cost.

Compare Top Models on WebArena-Verified

Last updated: August 29, 2026 · BenchLM version WebArena-Verified 2025

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.