# Scale Labs SciPredict (SciPredict)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-scipredict

- Category: [Knowledge](/knowledge)
- Last updated: September 29, 2026 snapshot

## About SciPredict

- Year: 2026
- Tasks: 15 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/scipredict)

BenchLM mirrors 15 published rows from the SciPredict public table captured on September 29, 2026 snapshot.

SciPredict is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (15 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [gemini-3-pro-preview](https://labs.scale.com/leaderboard/scipredict) | Google | 25.27% |
| 2 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 23.05% |
| 3 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 22.55% |
| 4 | [claude-opus-4-1-20250805](https://labs.scale.com/leaderboard/scipredict) | Anthropic | 22.22% |
| 5 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 22.22% |
| 6 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 20.58% |
| 7 | [o3-mini](/models/o3-mini) | OpenAI | 19.84% |
| 8 | [DeepSeek V3](/models/deepseek-v3) | DeepSeek | 19.18% |
| 9 | [Llama-3.3-70b](https://labs.scale.com/leaderboard/scipredict) | Meta | 18.19% |
| 10 | [o3](/models/o3) | OpenAI | 17.94% |
| 11 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 17.04% |
| 12 | [Qwen3-32B](https://labs.scale.com/leaderboard/scipredict) | Alibaba | 17.04% |
| 13 | [Qwen3-235B](https://labs.scale.com/leaderboard/scipredict) | Alibaba | 16.63% |
| 14 | [o4-mini-2025-04-03](https://labs.scale.com/leaderboard/scipredict) | OpenAI | 16.21% |
| 15 | [Llama-3.1-8B](https://labs.scale.com/leaderboard/scipredict) | Meta | 14.65% |

## FAQ

### What does SciPredict measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published SciPredict snapshot?

gemini-3-pro-preview currently leads the published SciPredict snapshot with a score of 25.27%.

### How many models are evaluated on SciPredict?

The September 29, 2026 snapshot contains 15 AI models.

### Does SciPredict affect BenchLM's overall score?

Not directly. SciPredict is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
