# FinanceArena — FinanceQA Assumption-Based (FinanceArena)

> An AfterQuery benchmark of open-ended financial analysis that requires models to read financial data, make assumptions, and return exact answers.

Canonical page: https://benchlm.ai/benchmarks/financearena

- Category: [Knowledge](/knowledge)
- Last updated: September 23, 2026 snapshot

## About FinanceArena

- Year: 2025
- Tasks: Professional financial-analysis questions
- Format: Open-ended financial QA with exact-match grading
- Difficulty: Professional finance reasoning
- Paper: [FinanceQA](https://arxiv.org/abs/2501.18062)

The mirrored AfterQuery table is labeled FinanceQA, Assumption-Based and reports exact-match accuracy. We keep it display-only because the embedded leaderboard rows do not expose the task count or per-row run configuration needed to merge them with other finance evaluations.

FinanceArena is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (19 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [o3](/models/o3) | OpenAI | 54.1% |
| 2 | [Grok 4](/models/grok-4) | xAI | 49.3% |
| 3 | [o4-mini (high)](/models/o4-mini-high) | OpenAI | 48.6% |
| 4 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 45.3% |
| 5 | [Llama 4 Maverick](/models/llama-4-maverick) | Meta | 44.6% |
| 6 | [Claude Opus 4](https://www.afterquery.com/leaderboard/finance-arena) | Anthropic | 44.6% |
| 7 | [Grok 3 [Beta]](/models/grok-3-beta) | xAI | 44.6% |
| 8 | [Claude Sonnet 4](https://www.afterquery.com/leaderboard/finance-arena) | Anthropic | 43.9% |
| 9 | [Phi-4 Reasoning Plus](https://www.afterquery.com/leaderboard/finance-arena) | Microsoft | 43.2% |
| 10 | [DeepSeek-R1](/models/deepseek-r1) | DeepSeek | 42.9% |
| 11 | [QwQ-32B](https://www.afterquery.com/leaderboard/finance-arena) | Alibaba | 42.6% |
| 12 | [GPT-4.1 mini](/models/gpt-4-1-mini) | OpenAI | 41.9% |
| 13 | [Llama 3.1 Nemotron Ultra](https://www.afterquery.com/leaderboard/finance-arena) | NVIDIA | 41.9% |
| 14 | [Qwen3 30B A3B](https://www.afterquery.com/leaderboard/finance-arena) | Alibaba | 37.2% |
| 15 | [Kimi K2](/models/kimi-k2) | Moonshot AI | 33.8% |
| 16 | [Gemini 2.5 Flash](/models/gemini-2-5-flash) | Google | 32.4% |
| 17 | [Magistral Medium](https://www.afterquery.com/leaderboard/finance-arena) | Mistral | 31.8% |
| 18 | [Command A](https://www.afterquery.com/leaderboard/finance-arena) | Cohere | 27.7% |
| 19 | [Nova Pro](/models/nova-pro) | Amazon | 20.3% |

## FAQ

### What does FinanceArena measure?

An AfterQuery benchmark of open-ended financial analysis that requires models to read financial data, make assumptions, and return exact answers.

### Which model leads the published FinanceArena snapshot?

o3 currently leads the published FinanceArena snapshot with a score of 54.1%.

### How many models are evaluated on FinanceArena?

The September 23, 2026 snapshot contains 19 AI models.

### Does FinanceArena affect BenchLM's overall score?

Not directly. FinanceArena is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
