Skip to main content
BenchLM

OfficeQA Pro

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A benchmark for grounded reasoning over office-style documents, spreadsheets, charts, and business artifacts.

Top models on OfficeQA Pro — September 27, 2026

As of September 27, 2026, Claude Opus 5.5 leads the OfficeQA Pro leaderboard with 67.7% , followed by Claude Opus 5 (66.9%) and Claude Opus 4.8 (66.2%).

12 modelsMultimodal & Grounded25% of category scoreCurrentUpdated September 27, 2026

Leaderboard (12 models)

Score
1
Claude Opus 5.5Anthropic · Closed
67.7%
2
Claude Opus 5Anthropic · Closed
66.9%
3
Claude Opus 4.8Anthropic · Closed
66.2%
4
Hy4 previewTencent · Open weight
66.2%
5
Kimi K3Moonshot AI · Closed
63.3%
6
GLM-5.3-FlashZ.AI · Open weight
62.4%
7
Step 5 PreviewStepFun · Closed
60.3%
8
Claude Fable 5Anthropic · Closed
57.9%
9
GPT-5.5OpenAI · Closed
54.1%
10
GPT-5.4OpenAI · Closed
53.2%
11
MiniMax M3MiniMax · Open weight
45.1%
12
Claude Opus 4.7 (Adaptive)Anthropic · Closed
43.6%

According to BenchLM.ai, Claude Opus 5.5 leads the OfficeQA Pro benchmark with a score of 67.7%, followed by Claude Opus 5 (66.9%) and Claude Opus 4.8 (66.2%). The top models are clustered within 1.5 points, suggesting this benchmark is nearing saturation for frontier models.

12 models have been evaluated on OfficeQA Pro. The benchmark falls in the Multimodal & Grounded category. The Multimodal & Grounded leaderboard ranks models by a weighted category score, and OfficeQA Pro contributes 25% of it. It does not enter the overall BenchAlign v5.7 ranking.

About OfficeQA Pro

Year

2026

Tasks

Document and spreadsheet tasks

Format

Grounded QA over office artifacts

Difficulty

Enterprise grounded reasoning

OfficeQA Pro is useful when choosing models for enterprise copilots because it measures whether they can reason correctly over real office content rather than generic chat prompts.

Freshness and provenance

Version

OfficeQA Pro 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does OfficeQA Pro measure?

A benchmark for grounded reasoning over office-style documents, spreadsheets, charts, and business artifacts.

Which model scores highest on OfficeQA Pro?

Claude Opus 5.5 by Anthropic currently leads with a score of 67.7% on OfficeQA Pro.

How many models are evaluated on OfficeQA Pro?

12 AI models have been evaluated on OfficeQA Pro on BenchLM.

Last updated: September 27, 2026 · BenchLM version OfficeQA Pro 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.