# Multi-Agent BrowseComp — 10-agent team prerelease configuration (BrowseComp (10-agent, prerelease))

> BrowseComp accuracy from ten collaborating Opus 5 agents on a pre-release model and unreleased effort configuration.

Canonical page: https://benchlm.ai/benchmarks/multiagentbrowsecompprerelease

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About BrowseComp (10-agent, prerelease)

- Year: 2026
- Tasks: BrowseComp web-research tasks
- Format: 10-agent team accuracy
- Difficulty: Long-horizon web research
- Paper: [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf)

Sections 8.11.1 and 8.11.4 report 93.6% for the 10-agent team. The card says these runs used a pre-release configuration without safeguards and are useful for relative, not absolute, comparison; BenchLM therefore keeps the row display-only and explicitly labeled.

BrowseComp (10-agent, prerelease) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 93.6% |

## FAQ

### What does BrowseComp (10-agent, prerelease) measure?

BrowseComp accuracy from ten collaborating Opus 5 agents on a pre-release model and unreleased effort configuration.

### Which model scores highest on BrowseComp (10-agent, prerelease)?

Claude Opus 5 by Anthropic currently leads with a score of 93.6% on BrowseComp (10-agent, prerelease).

### How many models are evaluated on BrowseComp (10-agent, prerelease)?

1 AI models have been evaluated on BrowseComp (10-agent, prerelease) on BenchLM.

### Does BrowseComp (10-agent, prerelease) affect BenchLM's overall score?

Not directly. BrowseComp (10-agent, prerelease) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
