# OfficeQA Pro

> A benchmark for grounded reasoning over office-style documents, spreadsheets, charts, and business artifacts.

Canonical page: https://benchlm.ai/benchmarks/officeqapro

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About OfficeQA Pro

- Year: 2026
- Tasks: Document and spreadsheet tasks
- Format: Grounded QA over office artifacts
- Difficulty: Enterprise grounded reasoning
- Paper: [OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning](https://arxiv.org/abs/2603.08655)

OfficeQA Pro is useful when choosing models for enterprise copilots because it measures whether they can reason correctly over real office content rather than generic chat prompts.

The Multimodal & Grounded leaderboard ranks models by a weighted category score, and OfficeQA Pro contributes 25% of it. It does not enter the overall BenchAlign v5.7 ranking.

## Leaderboard (12 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 67.7% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 66.9% |
| 3 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 66.2% |
| 4 | [Hy4 preview](/models/hy4-preview) | Tencent | 66.2% |
| 5 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 63.3% |
| 6 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 62.4% |
| 7 | [Step 5 Preview](/models/step-5-preview) | StepFun | 60.3% |
| 8 | [Claude Fable 5](/models/claude-fable) | Anthropic | 57.9% |
| 9 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 54.1% |
| 10 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 53.2% |
| 11 | [MiniMax M3](/models/minimax-m3) | MiniMax | 45.1% |
| 12 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 43.6% |

## FAQ

### What does OfficeQA Pro measure?

A benchmark for grounded reasoning over office-style documents, spreadsheets, charts, and business artifacts.

### Which model scores highest on OfficeQA Pro?

Claude Opus 5.5 by Anthropic currently leads with a score of 67.7% on OfficeQA Pro.

### How many models are evaluated on OfficeQA Pro?

12 AI models have been evaluated on OfficeQA Pro on BenchLM.

### Does OfficeQA Pro affect BenchLM's overall score?

Not the overall score. The Multimodal & Grounded leaderboard ranks models by a weighted category score, and OfficeQA Pro contributes 25% of it. It does not enter the overall BenchAlign v5.7 ranking.

## Compare Top Models on OfficeQA Pro

- [Claude Opus 5.5 vs Claude Opus 5](/compare/claude-opus-5-vs-claude-opus-5-5)
- [Claude Opus 5 vs Claude Opus 4.8](/compare/claude-opus-4-8-vs-claude-opus-5)
- [Claude Opus 4.8 vs Hy4 preview](/compare/claude-opus-4-8-vs-hy4-preview)
- [Hy4 preview vs Kimi K3](/compare/hy4-preview-vs-kimi-k3)
