WebVoyager
We mirror this table; we do not rank on it.
A browser-agent benchmark for completing multi-step workflows on live websites.
Benchmark score on WebVoyager — September 21, 2026
We mirror the published score view for WebVoyager. Fara1.5-27B leads the public snapshot at 89.3%, followed by Fara1.5-4B (80.8%). We do not use these results to rank models overall.
2 modelsAgenticCurrentDisplay onlyUpdated September 21, 2026
Benchmark score table (2 models)
ScoreAbout WebVoyager
Year
2026
Tasks
Live website workflows
Format
Interactive browser-agent evaluation
Difficulty
Multi-step web navigation
BenchLM stores WebVoyager as a display-only browser-agent benchmark reference outside the weighted ranking schema.
Freshness and provenance
Version
WebVoyager 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does WebVoyager measure?
A browser-agent benchmark for completing multi-step workflows on live websites.
Which model scores highest on WebVoyager?
Fara1.5-27B by Microsoft currently leads with a score of 89.3% on WebVoyager.
How many models are evaluated on WebVoyager?
2 AI models have been evaluated on WebVoyager on BenchLM.
Compare top models on WebVoyager
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.