Model profile
GLM-5
Evidence coverage
49 of 321 tracked benchmarks are published. 8 are verified and 41 provisional. 8 of 8 categories are measured.
- Published / tracked
- 49 / 321
- Verified
- 8
- Provisional
- 41
- Categories with evidence
- 8 / 8
Evidence by category
- Agentic13 benchmarksMixed evidence
- Coding7 benchmarksMixed evidence
- Reasoning4 benchmarksReported
- Knowledge12 benchmarksReported
- Math8 benchmarksMixed evidence
- Multilingual2 benchmarksReported
- Multimodal1 benchmarkReported
- Inst. Following2 benchmarksReported
GLM-5 ranks #28 out of 200 models on the public leaderboard with an overall score of 66.06/100. It also ranks #24 out of 99 on the verified leaderboard. This places it in the mid-tier of AI models, with strengths in specific benchmark categories.
GLM-5 is a open weight model with a 200K token context window. It processes queries without explicit chain-of-thought reasoning, offering faster response times and lower token usage.
GLM-5 sits inside the GLM-5 family alongside GLM-5.2, GLM-5.1, GLM-5 (Reasoning), GLM-5-Turbo, GLM-5V-Turbo. This profile currently has 49 of 321 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Its strongest category is Multilingual (#6), while its weakest is Coding (#24). This performance profile makes it a well-rounded choice across a range of tasks.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 65.2–66.89
- GLM-5-TurboZ.AICompare#2466.89GLM-5-Turbo is #24 with a score of 66.89.
- GPT-5.4 nanoOpenAICompare#2566.79GPT-5.4 nano is #25 with a score of 66.79.
- GPT-5.3 CodexOpenAICompare#2666.69GPT-5.3 Codex is #26 with a score of 66.69.
- Claude Opus 4.7 (Adaptive)AnthropicCompare#2766.27Claude Opus 4.7 (Adaptive) is #27 with a score of 66.27.
- GLM-5Current modelZ.AI#2866.06GLM-5 is #28 with a score of 66.06.
- Claude Sonnet 5AnthropicCompare#2965.32Claude Sonnet 5 is #29 with a score of 65.32.
- Qwen3.6 PlusAlibabaCompare#3065.2Qwen3.6 Plus is #30 with a score of 65.2.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
- Multilingual58%Eligible cohort rank #6 of 13Category score 48.7
- Math0%Eligible cohort rank #7 of 7Category score 57.2
- Inst. Following68%Eligible cohort rank #13 of 38Category score 86.5
- Knowledge75%Eligible cohort rank #14 of 52Category score 80.8
- Agentic81%Eligible cohort rank #23 of 119Category score 54.8
- Coding81%Eligible cohort rank #24 of 122Category score 59.4
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticRank #23 of 119Percentile 81stWeight 22%13 benchmarksMixed sources | 54.8 | #23 of 119 | 81st | 22% | 13 benchmarks | Mixed sources |
| CodingRank #24 of 122Percentile 81stWeight 20%7 benchmarksMixed sources | 59.4 | #24 of 122 | 81st | 20% | 7 benchmarks | Mixed sources |
| ReasoningRank Not rankedWeight 17%4 benchmarksReported | 78.9 | Not ranked | Not available | 17% | 4 benchmarks | Reported |
| KnowledgeRank #14 of 52Percentile 75thWeight 12%12 benchmarksReported | 80.8 | #14 of 52 | 75th | 12% | 12 benchmarks | Reported |
| MathRank #7 of 7Percentile 0thWeight 5%8 benchmarksMixed sources | 57.2 | #7 of 7 | 0th | 5% | 8 benchmarks | Mixed sources |
| MultilingualRank #6 of 13Percentile 58thWeight 7%2 benchmarksReported | 48.7 | #6 of 13 | 58th | 7% | 2 benchmarks | Reported |
| MultimodalRank Not rankedWeight 12%1 benchmarkReported | 68.8 | Not ranked | Not available | 12% | 1 benchmark | Reported |
| Inst. FollowingRank #13 of 38Percentile 68thWeight 5%2 benchmarksReported | 86.5 | #13 of 38 | 68th | 5% | 2 benchmarks | Reported |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1457 | ±4.4 | 27,814 |
| Coding | 1498 | ±7.5 | 7,210 |
| Math | 1443 | ±14.5 | 1,603 |
| Instruction Following | 1447 | ±6.8 | 8,910 |
| Creative Writing | 1445 | ±9.3 | 4,519 |
| Multi-turn | 1473 | ±9.2 | 4,437 |
| Hard Prompts | 1478 | ±5.3 | 17,146 |
| Hard Prompts (English) | 1487 | ±7.0 | 8,430 |
| Longer Query | 1470 | ±6.6 | 10,192 |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Agentic13 benchmarks
τ³-Bench Tool-Agent-User Evaluation
τ²-Bench Tool-Agent-User Evaluation
Gert Labs Composite Game Benchmark
Coding7 benchmarks
Software Engineering Benchmark Verified
SWE-bench Verified (mini-swe-agent-v2)
Artificial Analysis SciCode
Reasoning4 benchmarks
Artificial Analysis Long Context Reasoning
Critical Physics Tasks
Knowledge12 benchmarks
Humanity's Last Exam
Massive Multitask Language Understanding Professional
Graduate-Level Google-Proof Q&A
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines
GPQA Diamond
MMLU-Pro first-party comparison snapshot
Artificial Analysis GPQA Diamond
Artificial Analysis Humanity's Last Exam
Artificial Analysis Omniscience Index
Artificial Analysis Omniscience Accuracy
Artificial Analysis Omniscience Hallucination Rate
Math8 benchmarks
FrontierMath v2 Tiers 1-3
AIME 2026
Harvard-MIT Mathematics Tournament February 2026
FrontierMath v2 Tier 4
AIME25 first-party comparison snapshot
Harvard-MIT Mathematics Tournament February 2025
Harvard-MIT Mathematics Tournament November 2025
Multilingual2 benchmarks
Multimodal1 benchmark
Design Arena Website Elo
Inst. Following2 benchmarks
Instruction-Following Eval
Artificial Analysis IFBench
Frequently Asked Questions
How does GLM-5 perform overall in AI benchmarks?
GLM-5 currently ranks #28 out of 200 models on BenchLM's provisional leaderboard with an overall score of 66.06. It also ranks #24 out of 99 on the verified leaderboard. It is created by Z.AI. Its published context window is 200K.
Is GLM-5 good for knowledge and understanding?
GLM-5 ranks #14 out of 52 models in knowledge and understanding benchmarks with an average score of 80.8. There are stronger options in this category.
Is GLM-5 good for coding and programming?
GLM-5 ranks #24 out of 122 models in coding and programming benchmarks with an average score of 59.4. There are stronger options in this category.
Is GLM-5 good for mathematics?
GLM-5 ranks #7 out of 7 models in mathematics benchmarks with an average score of 57.2. It is among the top performers in this category.
Is GLM-5 good for reasoning and logic?
GLM-5 has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.
Is GLM-5 good for agentic tool use and computer tasks?
GLM-5 ranks #23 out of 119 models in agentic tool use and computer tasks benchmarks with an average score of 54.8. There are stronger options in this category.
Is GLM-5 good for multimodal and grounded tasks?
GLM-5 has visible benchmark coverage in multimodal and grounded tasks, but BenchLM does not currently assign it a global category rank there.
Is GLM-5 good for instruction following?
GLM-5 ranks #13 out of 38 models in instruction following benchmarks with an average score of 86.5. There are stronger options in this category.
Is GLM-5 good for multilingual tasks?
GLM-5 ranks #6 out of 13 models in multilingual tasks benchmarks with an average score of 48.7. It is among the top performers in this category.
Is GLM-5 open source?
Yes, GLM-5 is an open weight model created by Z.AI, meaning it can be downloaded and run locally or fine-tuned for specific use cases.
Which sibling models are related to GLM-5?
GLM-5 belongs to the GLM-5 family. Related variants on BenchLM include GLM-5.2, GLM-5.1, GLM-5 (Reasoning), GLM-5-Turbo, GLM-5V-Turbo.
Does GLM-5 have full benchmark coverage on BenchLM?
Not yet. GLM-5 currently has 49 published benchmark scores out of the 321 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of GLM-5?
GLM-5 has a published context window of 200K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.