Model profile
MiMo-V2.5
Evidence coverage
11 of 323 tracked benchmarks are published. 10 are verified and 1 provisional. 3 of 8 categories are measured.
- Published / tracked
- 11 / 323
- Verified
- 10
- Provisional
- 1
- Categories with evidence
- 3 / 8
Evidence by category
- Agentic5 benchmarksVerified
- Coding2 benchmarksVerified
- Reasoning0 benchmarksNot measured
- Knowledge0 benchmarksNot measured
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal4 benchmarksMixed evidence
- Inst. Following0 benchmarksNot measured
MiMo-V2.5 ranks #62 out of 200 models on the public leaderboard with an overall score of 58.62/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
MiMo-V2.5 is a proprietary model with a 1M token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
MiMo-V2.5 sits inside the MiMo-V2.5 family alongside MiMo-V2.5-Pro. BenchLM links it directly to MiMo-V2-Omni as the earlier related model in that lineage. This profile currently has 11 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Its strongest category is Multimodal & Grounded (#19), while its weakest is Coding (#60). This performance profile makes it particularly strong for screenshots, documents, charts, and grounded multimodal workflows.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 58.15–59.0
- GPT-5.2 InstantOpenAICompare#5959.0GPT-5.2 Instant is #59 with a score of 59.0.
- GPT-5.3 InstantOpenAICompare#6058.9GPT-5.3 Instant is #60 with a score of 58.9.
- DeepSeek V4 FlashDeepSeekCompare#6158.88DeepSeek V4 Flash is #61 with a score of 58.88.
- MiMo-V2.5Current modelXiaomi#6258.62MiMo-V2.5 is #62 with a score of 58.62.
- GPT-5 (high)OpenAICompare#6358.61GPT-5 (high) is #63 with a score of 58.61.
- GPT-5.2OpenAICompare#6458.43GPT-5.2 is #64 with a score of 58.43.
- DeepSeek V3.2 (Thinking)DeepSeekCompare#6558.15DeepSeek V3.2 (Thinking) is #65 with a score of 58.15.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
- Multimodal36%Eligible cohort rank #19 of 29Category score 63.3
- Agentic77%Eligible cohort rank #28 of 119Category score 53.8
- Coding51%Eligible cohort rank #60 of 122Category score 50.6
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticRank #28 of 119Percentile 77thWeight 22%5 benchmarksVerified | 53.8 | #28 of 119 | 77th | 22% | 5 benchmarks | Verified |
| CodingRank #60 of 122Percentile 51stWeight 20%2 benchmarksVerified | 50.6 | #60 of 122 | 51st | 20% | 2 benchmarks | Verified |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalRank #19 of 29Percentile 36thWeight 12%4 benchmarksMixed sources | 63.3 | #19 of 29 | 36th | 12% | 4 benchmarks | Mixed sources |
| Inst. FollowingWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1433 | ±4.5 | 42,281 |
| Coding | 1491 | ±6.9 | 12,219 |
| Math | 1440 | ±12.8 | 2,211 |
| Instruction Following | 1431 | ±6.5 | 14,034 |
| Creative Writing | 1392 | ±8.3 | 6,912 |
| Multi-turn | 1449 | ±8.0 | 7,452 |
| Hard Prompts | 1462 | ±5.2 | 28,157 |
| Hard Prompts (English) | 1471 | ±6.6 | 13,584 |
| Longer Query | 1452 | ±6.2 | 18,486 |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Agentic5 benchmarks
Gert Labs Composite Game Benchmark
Coding2 benchmarks
Multimodal4 benchmarks
Massive Multi-discipline Multimodal Understanding Pro
CharXiv Reasoning
Video-MME with subtitle
Design Arena Website Elo
Frequently Asked Questions
How does MiMo-V2.5 perform overall in AI benchmarks?
MiMo-V2.5 has 11 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is MiMo-V2.5 good for coding and programming?
MiMo-V2.5 ranks #60 out of 122 models in coding and programming benchmarks with an average score of 50.6. There are stronger options in this category.
Is MiMo-V2.5 good for agentic tool use and computer tasks?
MiMo-V2.5 ranks #28 out of 119 models in agentic tool use and computer tasks benchmarks with an average score of 53.8. There are stronger options in this category.
Is MiMo-V2.5 good for multimodal and grounded tasks?
MiMo-V2.5 ranks #19 out of 29 models in multimodal and grounded tasks benchmarks with an average score of 63.3. There are stronger options in this category.
Which sibling models are related to MiMo-V2.5?
MiMo-V2.5 belongs to the MiMo-V2.5 family. Related variants on BenchLM include MiMo-V2.5-Pro.
Does MiMo-V2.5 have full benchmark coverage on BenchLM?
Not yet. MiMo-V2.5 currently has 11 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of MiMo-V2.5?
MiMo-V2.5 has a published context window of 1M, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.