Benchmark profile
BrowseComp-VL
A vision-language browsing benchmark for multimodal web research and tool-use workflows.
Data verifiedHow BenchLM shows BrowseComp-VL right now
BenchLM is tracking BrowseComp-VL in the local dataset, but exact-source verification records for these rows are still being attached. To avoid a blank benchmark page, BenchLM shows the current tracked rows below as a display-only reference table.
These tracked rows are useful for inspection and spot-checking, but until exact-source attachments are completed they should not be treated as fully verified public benchmark rows.
Tracked score on BrowseComp-VL — July 29, 2026
BenchLM mirrors the published tracked score view for BrowseComp-VL. GLM-5V-Turbo leads the public snapshot at 51.9% , followed by Kimi K2.5 (42.9%) and Claude Opus 4.6 (35.9%). BenchLM does not use these results to rank models overall.
GLM-5V-Turbo
Z.AI
glm-5v-turbo
Kimi K2.5
Moonshot AI
kimi-k2-5
Claude Opus 4.6
Anthropic
claude-opus-4-6
Tracked score table (3 models)
ScoreThe published BrowseComp-VL snapshot places GLM-5V-Turbo first at 51.9%. The third row is 16.0 points behind. The broader top-10 range is 16.0 points, so the table still separates the published systems.
3 models have been evaluated on BrowseComp-VL. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. BrowseComp-VL is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About BrowseComp-VL
Year
2026
Tasks
Multimodal browsing tasks
Format
Vision-language web research evaluation
Difficulty
Multimodal browser-agent
BenchLM stores BrowseComp-VL as a display-only provider-table reference while keeping BrowseComp as the weighted core browsing benchmark.
BenchLM freshness & provenance
Version
BrowseComp-VL 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does BrowseComp-VL measure?
A vision-language browsing benchmark for multimodal web research and tool-use workflows.
Which model leads the published BrowseComp-VL snapshot?
GLM-5V-Turbo currently leads the published BrowseComp-VL snapshot with 51.9% tracked score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on BrowseComp-VL?
3 AI models are included in BenchLM's mirrored BrowseComp-VL snapshot, based on the public leaderboard captured on July 29, 2026.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.