Model profile
o1-preview
Evidence coverage
2 of 321 tracked benchmarks are published. 0 are verified and 2 provisional. 2 of 8 categories are measured.
- Published / tracked
- 2 / 321
- Verified
- 0
- Provisional
- 2
- Categories with evidence
- 2 / 8
Evidence by category
- Agentic0 benchmarksNot measured
- Coding1 benchmarkReported
- Reasoning0 benchmarksNot measured
- Knowledge1 benchmarkReported
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal0 benchmarksNot measured
- Inst. Following0 benchmarksNot measured
o1-preview ranks #123 out of 200 models on the public leaderboard with an overall score of 49.11/100. It also ranks #64 out of 99 on the verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
o1-preview is a proprietary model with a 200K token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
o1-preview sits inside the o1 family alongside o1, o1-pro. This profile currently has 2 of 321 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Its strongest category is Coding (#79). This performance profile makes it particularly well-suited for software development and code generation tasks.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 48.54–49.89
- DeepSeekMath V2DeepSeekCompare#11949.89DeepSeekMath V2 is #119 with a score of 49.89.
- Qwen2.5-1MAlibabaCompare#12049.89Qwen2.5-1M is #120 with a score of 49.89.
- Seed-2.0-LiteByteDanceCompare#12149.84Seed-2.0-Lite is #121 with a score of 49.84.
- Ministral 3 14B (Reasoning)MistralCompare#12249.34Ministral 3 14B (Reasoning) is #122 with a score of 49.34.
- o1-previewCurrent modelOpenAI#12349.11o1-preview is #123 with a score of 49.11.
- Aion-2.0Aion LabsCompare#12448.66Aion-2.0 is #124 with a score of 48.66.
- Mixtral 8x22B Instruct v0.1MistralCompare#12548.54Mixtral 8x22B Instruct v0.1 is #125 with a score of 48.54.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
- Coding36%Eligible cohort rank #79 of 122Category score 47.6
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%0 benchmarksNot measured | Not measured | Not ranked | Not available | 22% | 0 benchmarks | Not measured |
| CodingRank #79 of 122Percentile 36thWeight 20%1 benchmarkReported | 47.6 | #79 of 122 | 36th | 20% | 1 benchmark | Reported |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeRank Not rankedWeight 12%1 benchmarkReported | 60.5 | Not ranked | Not available | 12% | 1 benchmark | Reported |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| Inst. FollowingWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1388 | ±5.0 | 31,122 |
| Coding | 1416 | ±9.3 | 5,123 |
| Math | 1386 | ±9.7 | 4,569 |
| Instruction Following | 1381 | ±6.9 | 12,782 |
| Creative Writing | 1365 | ±9.9 | 4,508 |
| Multi-turn | 1390 | ±9.3 | 5,871 |
| Hard Prompts | 1396 | ±7.7 | 8,496 |
| Hard Prompts (English) | 1413 | ±9.5 | 4,917 |
| Longer Query | 1381 | ±9.7 | 4,578 |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Coding1 benchmark
Artificial Analysis Coding Index
Knowledge1 benchmark
Frequently Asked Questions
How does o1-preview perform overall in AI benchmarks?
o1-preview has 2 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is o1-preview good for knowledge and understanding?
o1-preview has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.
Is o1-preview good for coding and programming?
o1-preview ranks #79 out of 122 models in coding and programming benchmarks with an average score of 47.6. There are stronger options in this category.
Which sibling models are related to o1-preview?
o1-preview belongs to the o1 family. Related variants on BenchLM include o1, o1-pro.
Does o1-preview have full benchmark coverage on BenchLM?
Not yet. o1-preview currently has 2 published benchmark scores out of the 321 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of o1-preview?
o1-preview has a published context window of 200K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.