# OfficeQA

> Grounded numerical reasoning over a corpus of historical U.S. Treasury Bulletin documents.

Canonical page: https://benchlm.ai/benchmarks/officeqa

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About OfficeQA

- Year: 2026
- Tasks: Historical Treasury Bulletin questions
- Format: Agentic grounded QA accuracy
- Difficulty: Professional document reasoning
- Paper: [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf)

Section 8.13.1 reports the full-set result using extracted documents and code execution. Opus 5 ran through the public Messages API with production safeguards.

OfficeQA is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 78.9% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 78.1% |

## FAQ

### What does OfficeQA measure?

Grounded numerical reasoning over a corpus of historical U.S. Treasury Bulletin documents.

### Which model scores highest on OfficeQA?

Claude Opus 5.5 by Anthropic currently leads with a score of 78.9% on OfficeQA.

### How many models are evaluated on OfficeQA?

2 AI models have been evaluated on OfficeQA on BenchLM.

### Does OfficeQA affect BenchLM's overall score?

Not directly. OfficeQA is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on OfficeQA

- [Claude Opus 5.5 vs Claude Opus 5](/compare/claude-opus-5-vs-claude-opus-5-5)
