Skip to main content
BenchLM

Artificial Analysis AnalystAgent (AA-AnalystAgent)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Artificial Analysis' data-analysis benchmark, testing agents on spreadsheet and document work to answer the quantitative questions a business or data analyst faces day to day.

Benchmark score on AA-AnalystAgent — September 27, 2026

We compile the AA-AnalystAgent rows from secondary reports. Gemini 3.7 Flash leads the table at 60.0%, followed by Claude Fable 5.1 (57.5%) and Claude Opus 5 (53.8%). We do not use these results to rank models overall.

15 modelsAgenticCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (15 models)

Score
1
Gemini 3.7 FlashGoogle · Closed
60.0%
2
Claude Fable 5.1Anthropic · Closed
57.5%
3
Claude Opus 5Anthropic · Closed
53.8%
4
GPT-6 AstraOpenAI · Closed
51.2%
5
GPT-5.5OpenAI · Closed
50.0%
6
Claude Fable 5Anthropic · Closed
48.8%
7
GPT-5.6 SolOpenAI · Closed
47.5%
8
Claude Sonnet 5Anthropic · Closed
46.3%
9
Claude Opus 4.8Anthropic · Closed
45.0%
10
Grok 4.6xAI · Closed
41.3%
11
Kimi K3Moonshot AI · Closed
38.8%
12
InklingThinking Machines Lab · Open weight
23.8%
13
Mistral Medium 3.5 128BMistral · Open weight
12.5%
14
MiniMax M3MiniMax · Open weight
10.0%
15
Nemotron 3 UltraNVIDIA · Open weight
6.3%

Among the reported AA-AnalystAgent rows, Gemini 3.7 Flash is first at 60.0%. The third row is 6.2 points behind. The broader top-10 range is 18.7 points, so the table still separates the published systems.

15 models have been evaluated on AA-AnalystAgent. The benchmark falls in the Agentic category. AA-AnalystAgent is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-AnalystAgent

Year

2026

Tasks

Spreadsheet and document analysis questions

Format

Task success rate

Difficulty

Business and data analysis

Private dataset, independently run by Artificial Analysis. Display-only; not yet admitted as independent-run evidence.

Freshness and provenance

Version

AA-AnalystAgent 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does AA-AnalystAgent measure?

Artificial Analysis' data-analysis benchmark, testing agents on spreadsheet and document work to answer the quantitative questions a business or data analyst faces day to day.

Which model scores highest on AA-AnalystAgent?

Gemini 3.7 Flash by Google currently leads with a score of 60.0% on AA-AnalystAgent.

How many models are evaluated on AA-AnalystAgent?

15 AI models have been evaluated on AA-AnalystAgent on BenchLM.

Last updated: September 27, 2026 · BenchLM version AA-AnalystAgent 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.