Model profile
GPT-5 mini
Evidence coverage
1 of 321 tracked benchmarks is published. 1 is verified and 0 provisional. 1 of 8 categories are measured.
- Published / tracked
- 1 / 321
- Verified
- 1
- Provisional
- 0
- Categories with evidence
- 1 / 8
Evidence by category
- Agentic0 benchmarksNot measured
- Coding1 benchmarkVerified
- Reasoning0 benchmarksNot measured
- Knowledge0 benchmarksNot measured
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal0 benchmarksNot measured
- Inst. Following0 benchmarksNot measured
GPT-5 mini ranks #153 out of 200 models on the public leaderboard with an overall score of 43.93/100. It also ranks #73 out of 99 on the verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
GPT-5 mini is a proprietary model with a 128K token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
GPT-5 mini sits inside the GPT-5 family alongside GPT-5 (high), GPT-5 (medium), GPT-5 nano. This profile currently has 1 of 321 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Its strongest category is Coding (#105). This performance profile makes it particularly well-suited for software development and code generation tasks.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 43.2–44.6
- Seed-2.0-MiniByteDanceCompare#14944.6Seed-2.0-Mini is #149 with a score of 44.6.
- Nemotron Ultra 253BNVIDIACompare#15044.4Nemotron Ultra 253B is #150 with a score of 44.4.
- Nemotron 3 Nano Omni 30B A3BNVIDIACompare#15144.24Nemotron 3 Nano Omni 30B A3B is #151 with a score of 44.24.
- GPT-4.1 miniOpenAICompare#15244.19GPT-4.1 mini is #152 with a score of 44.19.
- GPT-5 miniCurrent modelOpenAI#15343.93GPT-5 mini is #153 with a score of 43.93.
- Ling 2.6 FlashInclusionAICompare#15443.87Ling 2.6 Flash is #154 with a score of 43.87.
- Gemma 4 E4BGoogleCompare#15543.2Gemma 4 E4B is #155 with a score of 43.2.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
- Coding14%Eligible cohort rank #105 of 122Category score 40.5
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%0 benchmarksNot measured | Not measured | Not ranked | Not available | 22% | 0 benchmarks | Not measured |
| CodingRank #105 of 122Percentile 14thWeight 20%1 benchmarkVerified | 40.5 | #105 of 122 | 14th | 20% | 1 benchmark | Verified |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| Inst. FollowingWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1243 | Not available | Not available |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Coding1 benchmark
Vibe Code Bench v1.1
Frequently Asked Questions
How does GPT-5 mini perform overall in AI benchmarks?
GPT-5 mini has 1 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is GPT-5 mini good for coding and programming?
GPT-5 mini ranks #105 out of 122 models in coding and programming benchmarks with an average score of 40.5. There are stronger options in this category.
Which sibling models are related to GPT-5 mini?
GPT-5 mini belongs to the GPT-5 family. Related variants on BenchLM include GPT-5 (high), GPT-5 (medium), GPT-5 nano.
Does GPT-5 mini have full benchmark coverage on BenchLM?
Not yet. GPT-5 mini currently has 1 published benchmark scores out of the 321 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of GPT-5 mini?
GPT-5 mini has a published context window of 128K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.