Skip to main content
BenchLM

Artificial Analysis GDP.pdf (GDP.pdf)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Artificial Analysis' document-output benchmark: professional tasks whose deliverable is a produced PDF, one of the ten components of its Intelligence Index v4.3.

Benchmark score on GDP.pdf — September 27, 2026

We compile the GDP.pdf rows from secondary reports. GPT-6 Astra leads the table at 31.0%, followed by GPT-5.6 Sol (27.2%) and Muse Spark 1.3 (26.6%). We do not use these results to rank models overall.

14 modelsAgenticCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (14 models)

Score
1
GPT-6 AstraOpenAI · Closed
31.0%
2
GPT-5.6 SolOpenAI · Closed
27.2%
3
Muse Spark 1.3Meta · Closed
26.6%
4
Claude Fable 5.1Anthropic · Closed
26.2%
5
Claude Opus 5.5Anthropic · Closed
26.2%
6
GPT-6 SolOpenAI · Closed
24.8%
7
GPT-5.6 LunaOpenAI · Closed
24.0%
8
Kimi K3Moonshot AI · Closed
22.0%
9
Claude Opus 5Anthropic · Closed
21.6%
10
Gemini 3.8 FlashGoogle · Closed
21.0%
11
GPT-6 LunaOpenAI · Closed
20.4%
12
Grok 4.7xAI · Closed
20.0%
13
MiMo-V2.6-ProXiaomi · Open weight
19.2%
14
Grok 4.6xAI · Closed
17.0%

Among the reported GDP.pdf rows, GPT-6 Astra is first at 31.0%. The third row is 4.4 points behind. The broader top-10 range is 10.0 points, so the table still separates the published systems.

14 models have been evaluated on GDP.pdf. The benchmark falls in the Agentic category. GDP.pdf is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About GDP.pdf

Year

2026

Tasks

Professional document-production tasks

Format

Task success rate

Difficulty

Professional knowledge work

Independently run by Artificial Analysis under its published harness. Display-only; not yet admitted as independent-run evidence.

Freshness and provenance

Version

GDP.pdf 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does GDP.pdf measure?

Artificial Analysis' document-output benchmark: professional tasks whose deliverable is a produced PDF, one of the ten components of its Intelligence Index v4.3.

Which model scores highest on GDP.pdf?

GPT-6 Astra by OpenAI currently leads with a score of 31.0% on GDP.pdf.

How many models are evaluated on GDP.pdf?

14 AI models have been evaluated on GDP.pdf on BenchLM.

Last updated: September 27, 2026 · BenchLM version GDP.pdf 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.