Model profile
GLM-4.5
Evidence coverage
1 of 321 tracked benchmarks is published. 0 are verified and 1 provisional. 1 of 8 categories are measured.
- Published / tracked
- 1 / 321
- Verified
- 0
- Provisional
- 1
- Categories with evidence
- 1 / 8
Evidence by category
- Agentic0 benchmarksNot measured
- Coding0 benchmarksNot measured
- Reasoning0 benchmarksNot measured
- Knowledge0 benchmarksNot measured
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal1 benchmarkReported
- Inst. Following0 benchmarksNot measured
GLM-4.5 ranks #68 out of 200 models on the public leaderboard with an overall score of 57.56/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
GLM-4.5 is a proprietary model with a 128K token context window. It processes queries without explicit chain-of-thought reasoning, offering faster response times and lower token usage.
This profile currently has 1 of 321 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 57.01–58.15
- DeepSeek V3.2 (Thinking)DeepSeekCompare#6558.15DeepSeek V3.2 (Thinking) is #65 with a score of 58.15.
- Qwen3 235B 2507 (Reasoning)AlibabaCompare#6658.01Qwen3 235B 2507 (Reasoning) is #66 with a score of 58.01.
- Gemma 4 26B A4BGoogleCompare#6757.96Gemma 4 26B A4B is #67 with a score of 57.96.
- GLM-4.5Current modelZ.AI#6857.56GLM-4.5 is #68 with a score of 57.56.
- Claude Opus 4.5 ThinkingAnthropicCompare#6957.44Claude Opus 4.5 Thinking is #69 with a score of 57.44.
- Gemini 2.5 ProGoogleCompare#7057.25Gemini 2.5 Pro is #70 with a score of 57.25.
- Qwen3.5 397BAlibabaCompare#7157.01Qwen3.5 397B is #71 with a score of 57.01.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%0 benchmarksNot measured | Not measured | Not ranked | Not available | 22% | 0 benchmarks | Not measured |
| CodingWeight 20%0 benchmarksNot measured | Not measured | Not ranked | Not available | 20% | 0 benchmarks | Not measured |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalRank Not rankedWeight 12%1 benchmarkReported | 40.4 | Not ranked | Not available | 12% | 1 benchmark | Reported |
| Inst. FollowingWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1411 | ±4.9 | 24,292 |
| Coding | 1454 | ±8.8 | 4,770 |
| Math | 1413 | ±15.5 | 1,425 |
| Instruction Following | 1405 | ±7.8 | 6,165 |
| Creative Writing | 1373 | ±10.9 | 3,127 |
| Multi-turn | 1406 | ±9.7 | 3,899 |
| Hard Prompts | 1433 | ±6.3 | 11,198 |
| Hard Prompts (English) | 1432 | ±8.1 | 5,692 |
| Longer Query | 1417 | ±8.6 | 5,011 |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Multimodal1 benchmark
Design Arena Website Elo
Frequently Asked Questions
How does GLM-4.5 perform overall in AI benchmarks?
GLM-4.5 has 1 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is GLM-4.5 good for multimodal and grounded tasks?
GLM-4.5 has visible benchmark coverage in multimodal and grounded tasks, but BenchLM does not currently assign it a global category rank there.
Does GLM-4.5 have full benchmark coverage on BenchLM?
Not yet. GLM-4.5 currently has 1 published benchmark scores out of the 321 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of GLM-4.5?
GLM-4.5 has a published context window of 128K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.