# ArXivMath August 2026 with tools (ArXivMath Aug. 2026 (tools))

> The August 2026 ArXivMath release evaluated with a code-execution sandbox and no internet access.

Canonical page: https://benchlm.ai/benchmarks/arxivmathaugust2026withtools

- Category: [Mathematics](/math)
- Last updated: September 27, 2026

## About ArXivMath Aug. 2026 (tools)

- Year: 2026
- Tasks: 57 recent research-mathematics problems
- Format: Final-answer accuracy with tools
- Difficulty: Research mathematics
- Paper: [Claude Opus 5.5 System Card](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf)

Section 8.9 reports four-attempt average accuracy on the distinct 57-problem August 2026 release at max effort with tools. It remains separate from June and from MathArena's vendor-harness leaderboard.

ArXivMath Aug. 2026 (tools) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 96.9% |

## FAQ

### What does ArXivMath Aug. 2026 (tools) measure?

The August 2026 ArXivMath release evaluated with a code-execution sandbox and no internet access.

### Which model scores highest on ArXivMath Aug. 2026 (tools)?

Claude Opus 5.5 by Anthropic currently leads with a score of 96.9% on ArXivMath Aug. 2026 (tools).

### How many models are evaluated on ArXivMath Aug. 2026 (tools)?

1 AI models have been evaluated on ArXivMath Aug. 2026 (tools) on BenchLM.

### Does ArXivMath Aug. 2026 (tools) affect BenchLM's overall score?

Not directly. ArXivMath Aug. 2026 (tools) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
