# Scale Labs PropensityBench (PropensityBench)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-propensitybench

- Category: [Reasoning](/reasoning)
- Last updated: September 29, 2026 snapshot

## About PropensityBench

- Year: 2026
- Tasks: 14 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/propensitybench)

BenchLM mirrors 14 published rows from the PropensityBench public table captured on September 29, 2026 snapshot.

PropensityBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (14 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 79.00% |
| 2 | [gemini-2.0-flash](https://labs.scale.com/leaderboard/propensitybench) | Google | 77.80% |
| 3 | [Qwen3-8B](https://labs.scale.com/leaderboard/propensitybench) | Alibaba | 75.20% |
| 4 | [\t\ngemini-2.5-flash](https://labs.scale.com/leaderboard/propensitybench) | Google | 68.00% |
| 5 | [Llama-3.1-8B](https://labs.scale.com/leaderboard/propensitybench) | Meta | 66.50% |
| 6 | [Llama-3.1-70B](https://labs.scale.com/leaderboard/propensitybench) | Meta | 55.40% |
| 7 | [gemini-3-pro-preview](https://labs.scale.com/leaderboard/propensitybench) | Google | 52.85% |
| 8 | [GPT-4o](/models/gpt-4o) | OpenAI | 46.10% |
| 9 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 34.35% |
| 10 | [o3-mini](/models/o3-mini) | OpenAI | 33.20% |
| 11 | [Qwen2.5-32B](https://labs.scale.com/leaderboard/propensitybench) | Alibaba | 22.90% |
| 12 | [o4-mini-2025-04-16](https://labs.scale.com/leaderboard/propensitybench) | OpenAI | 15.80% |
| 13 | [claude-sonnet-4-20250514](https://labs.scale.com/leaderboard/propensitybench) | Anthropic | 12.20% |
| 14 | [o3](/models/o3) | OpenAI | 10.50% |

## FAQ

### What does PropensityBench measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published PropensityBench snapshot?

Gemini 2.5 Pro currently leads the published PropensityBench snapshot with a score of 79.00%.

### How many models are evaluated on PropensityBench?

The September 29, 2026 snapshot contains 14 AI models.

### Does PropensityBench affect BenchLM's overall score?

Not directly. PropensityBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
