# ArXivMath June 2026 with tools (ArXivMath Jun. 2026 (tools))

> Final-answer research mathematics problems drawn from recent arXiv abstracts with tool access.

Canonical page: https://benchlm.ai/benchmarks/arxivmathjune2026withtools

- Category: [Mathematics](/math)
- Last updated: September 27, 2026

## About ArXivMath Jun. 2026 (tools)

- Year: 2026
- Tasks: 49 recent research-mathematics problems
- Format: Final-answer accuracy with tools
- Difficulty: Research mathematics
- Paper: [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf)

Section 8.8 reports four-run average accuracy on the 49-problem June 2026 release at max effort with tools.

ArXivMath Jun. 2026 (tools) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 91.3% |

## FAQ

### What does ArXivMath Jun. 2026 (tools) measure?

Final-answer research mathematics problems drawn from recent arXiv abstracts with tool access.

### Which model scores highest on ArXivMath Jun. 2026 (tools)?

Claude Opus 5 by Anthropic currently leads with a score of 91.3% on ArXivMath Jun. 2026 (tools).

### How many models are evaluated on ArXivMath Jun. 2026 (tools)?

1 AI models have been evaluated on ArXivMath Jun. 2026 (tools) on BenchLM.

### Does ArXivMath Jun. 2026 (tools) affect BenchLM's overall score?

Not directly. ArXivMath Jun. 2026 (tools) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
