# ArXivMath August 2026 without tools (ArXivMath Aug. 2026 (no tools))

> Final-answer research mathematics problems drawn from recent arXiv papers, including problems based on resolved or refuted conjectures.

Canonical page: https://benchlm.ai/benchmarks/arxivmathaugust2026

- Category: [Mathematics](/math)
- Last updated: September 27, 2026

## About ArXivMath Aug. 2026 (no tools)

- Year: 2026
- Tasks: 57 recent research-mathematics problems
- Format: Final-answer accuracy
- Difficulty: Research mathematics
- Paper: [Claude Opus 5.5 System Card](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf)

Section 8.9 reports four-attempt average accuracy on the distinct 57-problem August 2026 release at max effort without tools. It is not comparable to the June release and remains display-only.

ArXivMath Aug. 2026 (no tools) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 91.2% |

## FAQ

### What does ArXivMath Aug. 2026 (no tools) measure?

Final-answer research mathematics problems drawn from recent arXiv papers, including problems based on resolved or refuted conjectures.

### Which model scores highest on ArXivMath Aug. 2026 (no tools)?

Claude Opus 5.5 by Anthropic currently leads with a score of 91.2% on ArXivMath Aug. 2026 (no tools).

### How many models are evaluated on ArXivMath Aug. 2026 (no tools)?

1 AI models have been evaluated on ArXivMath Aug. 2026 (no tools) on BenchLM.

### Does ArXivMath Aug. 2026 (no tools) affect BenchLM's overall score?

Not directly. ArXivMath Aug. 2026 (no tools) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
