# Artificial Analysis AnalystAgent (AA-AnalystAgent)

> Artificial Analysis' data-analysis benchmark, testing agents on spreadsheet and document work to answer the quantitative questions a business or data analyst faces day to day.

Canonical page: https://benchlm.ai/benchmarks/aaanalystagent

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About AA-AnalystAgent

- Year: 2026
- Tasks: Spreadsheet and document analysis questions
- Format: Task success rate
- Difficulty: Business and data analysis
- Paper: [AA-AnalystAgent Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-analyst-agent)

Private dataset, independently run by Artificial Analysis. Display-only; not yet admitted as independent-run evidence.

AA-AnalystAgent is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (15 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 60.0% |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 57.5% |
| 3 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 53.8% |
| 4 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 51.2% |
| 5 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 50.0% |
| 6 | [Claude Fable 5](/models/claude-fable) | Anthropic | 48.8% |
| 7 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 47.5% |
| 8 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 46.3% |
| 9 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 45.0% |
| 10 | [Grok 4.6](/models/grok-4-6) | xAI | 41.3% |
| 11 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 38.8% |
| 12 | [Inkling](/models/inkling) | Thinking Machines Lab | 23.8% |
| 13 | [Mistral Medium 3.5 128B](/models/mistral-medium-3-5-128b) | Mistral | 12.5% |
| 14 | [MiniMax M3](/models/minimax-m3) | MiniMax | 10.0% |
| 15 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 6.3% |

## FAQ

### What does AA-AnalystAgent measure?

Artificial Analysis' data-analysis benchmark, testing agents on spreadsheet and document work to answer the quantitative questions a business or data analyst faces day to day.

### Which model scores highest on AA-AnalystAgent?

Gemini 3.7 Flash by Google currently leads with a score of 60.0% on AA-AnalystAgent.

### How many models are evaluated on AA-AnalystAgent?

15 AI models have been evaluated on AA-AnalystAgent on BenchLM.

### Does AA-AnalystAgent affect BenchLM's overall score?

Not directly. AA-AnalystAgent is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on AA-AnalystAgent

- [Gemini 3.7 Flash vs Claude Fable 5.1](/compare/claude-fable-5-1-vs-gemini-3-7-flash)
- [Claude Fable 5.1 vs Claude Opus 5](/compare/claude-fable-5-1-vs-claude-opus-5)
- [Claude Opus 5 vs GPT-6 Astra](/compare/claude-opus-5-vs-gpt-6-astra)
- [GPT-6 Astra vs GPT-5.5](/compare/gpt-5-5-vs-gpt-6-astra)
