Model profile
o1-pro
Evidence coverage
2 of 323 tracked benchmarks are published. 0 are verified and 2 provisional. 1 of 8 categories are measured.
- Published / tracked
- 2 / 323
- Verified
- 0
- Provisional
- 2
- Categories with evidence
- 1 / 8
Evidence by category
- Agentic0 benchmarksNot measured
- Coding0 benchmarksNot measured
- Reasoning0 benchmarksNot measured
- Knowledge2 benchmarksReported
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal0 benchmarksNot measured
- Inst. Following0 benchmarksNot measured
o1-pro ranks #141 out of 200 models on the public leaderboard with an overall score of 45.94/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
o1-pro is a proprietary model with a 200K token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
o1-pro sits inside the o1 family alongside o1, o1-preview. This profile currently has 2 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 45.08–46.36
- GPT-5 nanoOpenAICompare#13946.36GPT-5 nano is #139 with a score of 46.36.
- Mistral Small 4MistralCompare#14046.23Mistral Small 4 is #140 with a score of 46.23.
- o1-proCurrent modelOpenAI#14145.94o1-pro is #141 with a score of 45.94.
- Claude 4.1 OpusAnthropicCompare#14245.88Claude 4.1 Opus is #142 with a score of 45.88.
- Z-1ZCompare#14345.13Z-1 is #143 with a score of 45.13.
- Nemotron-4 15BNVIDIACompare#14445.08Nemotron-4 15B is #144 with a score of 45.08.
- Seed 1.6 FlashByteDanceCompare#14545.08Seed 1.6 Flash is #145 with a score of 45.08.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%0 benchmarksNot measured | Not measured | Not ranked | Not available | 22% | 0 benchmarks | Not measured |
| CodingWeight 20%0 benchmarksNot measured | Not measured | Not ranked | Not available | 20% | 0 benchmarks | Not measured |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeRank Not rankedWeight 12%2 benchmarksReported | 79.0 | Not ranked | Not available | 12% | 2 benchmarks | Reported |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| Inst. FollowingWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Knowledge2 benchmarks
Graduate-Level Google-Proof Q&A
Frequently Asked Questions
How does o1-pro perform overall in AI benchmarks?
o1-pro has 2 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is o1-pro good for knowledge and understanding?
o1-pro has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.
Which sibling models are related to o1-pro?
o1-pro belongs to the o1 family. Related variants on BenchLM include o1, o1-preview.
Does o1-pro have full benchmark coverage on BenchLM?
Not yet. o1-pro currently has 2 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of o1-pro?
o1-pro has a published context window of 200K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.