# Scale Labs PRBench Finance (PRBench Finance)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-prbench-finance

- Category: [Knowledge](/knowledge)
- Last updated: September 29, 2026 snapshot

## About PRBench Finance

- Year: 2026
- Tasks: 35 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/prbench-finance)

BenchLM mirrors 35 published rows from the PRBench Finance public table captured on September 29, 2026 snapshot.

PRBench Finance is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (35 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 59.54% |
| 2 | [Muse Spark 1.1\n](https://labs.scale.com/leaderboard/prbench-finance) | Meta | 55.01% |
| 3 | [claude-fable 5\n](https://labs.scale.com/leaderboard/prbench-finance) | Anthropic | 53.86% |
| 4 | [claude-opus-4-6 (Non-Thinking)](https://labs.scale.com/leaderboard/prbench-finance) | Anthropic | 53.28% |
| 5 | [Muse Spark](/models/muse-spark) | Meta | 52.44% |
| 6 | [gpt-5](https://labs.scale.com/leaderboard/prbench-finance) | OpenAI | 51.32% |
| 7 | [gpt-5-pro](https://labs.scale.com/leaderboard/prbench-finance) | OpenAI | 51.06% |
| 8 | [Fable 5.1\n](https://labs.scale.com/leaderboard/prbench-finance) | Anthropic | 50.80% |
| 9 | [gpt-5.6-sol (max)](https://labs.scale.com/leaderboard/prbench-finance) | OpenAI | 50.45% |
| 10 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 49.51% |
| 11 | [o3-pro](/models/o3-pro) | OpenAI | 49.08% |
| 12 | [gpt-5.1-thinking](https://labs.scale.com/leaderboard/prbench-finance) | OpenAI | 48.01% |
| 13 | [o3](/models/o3) | OpenAI | 47.69% |
| 14 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 47.54% |
| 15 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 46.51% |
| 16 | [GPT-5.2 Pro](/models/gpt-5-2-pro) | OpenAI | 46.34% |
| 17 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | Anthropic | 46.16% |
| 18 | [gpt-5.4 (High)](https://labs.scale.com/leaderboard/prbench-finance) | OpenAI | 45.63% |
| 19 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 43.80% |
| 20 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 43.79% |
| 21 | [kimi-k2-thinking](https://labs.scale.com/leaderboard/prbench-finance) | Moonshot AI | 43.41% |
| 22 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 41.87% |
| 23 | [mistral-medium-latest](https://labs.scale.com/leaderboard/prbench-finance) | Mistral | 39.35% |
| 24 | [o4-mini](https://labs.scale.com/leaderboard/prbench-finance) | OpenAI | 39.22% |
| 25 | [gemini-3-pro-preview](https://labs.scale.com/leaderboard/prbench-finance) | Google | 39.18% |
| 26 | [qwen.qwen3-235b-a22b-2507-v1:0](https://labs.scale.com/leaderboard/prbench-finance) | Alibaba | 39.14% |
| 27 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 38.92% |
| 28 | [Gemini 2.5 Flash](/models/gemini-2-5-flash) | Google | 38.41% |
| 29 | [kimi-k2-instruct](https://labs.scale.com/leaderboard/prbench-finance) | Moonshot AI | 38.34% |
| 30 | [claude-opus-4-1-20250805](https://labs.scale.com/leaderboard/prbench-finance) | Anthropic | 35.15% |
| 31 | [deepseek-v3p1](https://labs.scale.com/leaderboard/prbench-finance) | Deepseek | 35.09% |
| 32 | [GPT-4.1](/models/gpt-4-1) | OpenAI | 34.32% |
| 33 | [deepseek-r1-0528](https://labs.scale.com/leaderboard/prbench-finance) | Deepseek | 32.67% |
| 34 | [GPT-4.1 mini](/models/gpt-4-1-mini) | OpenAI | 30.45% |
| 35 | [llama4-maverick-instruct-basic](https://labs.scale.com/leaderboard/prbench-finance) | Meta | 22.36% |

## FAQ

### What does PRBench Finance measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published PRBench Finance snapshot?

Muse Spark 1.3 currently leads the published PRBench Finance snapshot with a score of 59.54%.

### How many models are evaluated on PRBench Finance?

The September 29, 2026 snapshot contains 35 AI models.

### Does PRBench Finance affect BenchLM's overall score?

Not directly. PRBench Finance is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
