Model profile
GLM-5 (Reasoning)
Evidence coverage
2 of 321 tracked benchmarks are published. 1 is verified and 1 provisional. 2 of 8 categories are measured.
- Published / tracked
- 2 / 321
- Verified
- 1
- Provisional
- 1
- Categories with evidence
- 2 / 8
Evidence by category
- Agentic0 benchmarksNot measured
- Coding1 benchmarkVerified
- Reasoning0 benchmarksNot measured
- Knowledge0 benchmarksNot measured
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal1 benchmarkReported
- Inst. Following0 benchmarksNot measured
GLM-5 (Reasoning) ranks #52 out of 200 models on the public leaderboard with an overall score of 59.77/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
GLM-5 (Reasoning) is a open weight model with a 200K token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
GLM-5 (Reasoning) sits inside the GLM-5 family alongside GLM-5, GLM-5.2, GLM-5.1, GLM-5-Turbo, GLM-5V-Turbo. This profile currently has 2 of 321 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 59.35–59.97
- Grok 4.1xAICompare#5159.97Grok 4.1 is #51 with a score of 59.97.
- GLM-5 (Reasoning)Current modelZ.AI#5259.77GLM-5 (Reasoning) is #52 with a score of 59.77.
- Qwen 3.6 Max (preview)AlibabaCompare#5359.72Qwen 3.6 Max (preview) is #53 with a score of 59.72.
- Kimi K2.5Moonshot AICompare#5459.66Kimi K2.5 is #54 with a score of 59.66.
- MiniMax M2.5MiniMaxCompare#5559.52MiniMax M2.5 is #55 with a score of 59.52.
- Qwen3.5 397B (Reasoning)AlibabaCompare#5659.5Qwen3.5 397B (Reasoning) is #56 with a score of 59.5.
- Kimi K2.5 (Reasoning)Moonshot AICompare#5759.35Kimi K2.5 (Reasoning) is #57 with a score of 59.35.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%0 benchmarksNot measured | Not measured | Not ranked | Not available | 22% | 0 benchmarks | Not measured |
| CodingRank Not rankedWeight 20%1 benchmarkVerified | 57.2 | Not ranked | Not available | 20% | 1 benchmark | Verified |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalRank Not rankedWeight 12%1 benchmarkReported | 78.0 | Not ranked | Not available | 12% | 1 benchmark | Reported |
| Inst. FollowingWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1456 | ±6.0 | 11,101 |
| Coding | 1491 | ±11.6 | 2,492 |
| Math | 1455 | ±21.0 | 717 |
| Instruction Following | 1445 | ±10.4 | 3,132 |
| Creative Writing | 1442 | ±14.6 | 1,725 |
| Multi-turn | 1457 | ±13.9 | 1,747 |
| Hard Prompts | 1476 | ±7.6 | 6,159 |
| Hard Prompts (English) | 1480 | ±10.9 | 2,816 |
| Longer Query | 1462 | ±10.4 | 3,140 |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Coding1 benchmark
Vibe Code Bench v1.1
Multimodal1 benchmark
Design Arena Website Elo
Frequently Asked Questions
How does GLM-5 (Reasoning) perform overall in AI benchmarks?
GLM-5 (Reasoning) has 2 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is GLM-5 (Reasoning) good for coding and programming?
GLM-5 (Reasoning) has visible benchmark coverage in coding and programming, but BenchLM does not currently assign it a global category rank there.
Is GLM-5 (Reasoning) good for multimodal and grounded tasks?
GLM-5 (Reasoning) has visible benchmark coverage in multimodal and grounded tasks, but BenchLM does not currently assign it a global category rank there.
Is GLM-5 (Reasoning) open source?
Yes, GLM-5 (Reasoning) is an open weight model created by Z.AI, meaning it can be downloaded and run locally or fine-tuned for specific use cases.
Which sibling models are related to GLM-5 (Reasoning)?
GLM-5 (Reasoning) belongs to the GLM-5 family. Related variants on BenchLM include GLM-5, GLM-5.2, GLM-5.1, GLM-5-Turbo, GLM-5V-Turbo.
Does GLM-5 (Reasoning) have full benchmark coverage on BenchLM?
Not yet. GLM-5 (Reasoning) currently has 2 published benchmark scores out of the 321 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of GLM-5 (Reasoning)?
GLM-5 (Reasoning) has a published context window of 200K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.