Model comparison
Claude Opus 4.5 vs GLM-5
Head-to-head evidence from 42 shared benchmark results across 8 categories. Overall scores shown here use the public BenchAlign v5 ranking lane.
Public leaderboard positions: Claude Opus 4.5 #34 (Supported); GLM-5 #28 (Supported). Intervals and evidence labels describe ranking uncertainty, not a guarantee for a specific workload.
Evidence parity. Claude Opus 4.5 and GLM-5 share 42 comparable benchmark results. 7 of 8 categories are comparable. 17 results are unique to Claude Opus 4.5; 7 to GLM-5.
Updated July 22, 2026- Shared results
- 42
- Claude Opus 4.5 only
- 17
- GLM-5 only
- 7
- Comparable categories
- 7 / 8
Pick GLM-5 if you want the stronger benchmark profile. Claude Opus 4.5 only becomes the better choice if agentic is the priority.
Confidence note. This is a partial-evidence comparison with 42 shared benchmark results across 8 evidence categories; 7 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.
Why this result
GLM-5 has the cleaner BenchAlign overall profile here, landing at 66.06 versus 64.22. It is a real lead, but still close enough that category-level strengths matter more than the headline number.
GLM-5's sharpest advantage is in instruction following, where it averages 92.6 against 69.5. The single biggest benchmark swing on the page is HLE, 30.8% to 50.4%. Claude Opus 4.5 does hit back in agentic, so the answer changes if that is the part of the workload you care about most.
Claude Opus 4.5 is also the more expensive model on tokens at $5.00 input / $25.00 output per 1M tokens, versus $1.00 input / $3.20 output per 1M tokens for GLM-5. That is roughly 7.8x on output cost alone.
Category breakdown
Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.
| Category | Claude Opus 4.5 | Δ | GLM-5 |
|---|---|---|---|
| Inst. Following | Claude Opus 4.569.5 | Margin→ 23.1 | GLM-592.6 |
| Knowledge | Claude Opus 4.558.1 | Margin→ 8.3 | GLM-566.4 |
| Agentic | Claude Opus 4.562.6 | Margin← 6.4 | GLM-556.2 |
| Coding | Claude Opus 4.571.7 | Margin← 5.4 | GLM-566.3 |
| Reasoning | Claude Opus 4.564.4 | Margin← 3.6 | GLM-560.8 |
| Multilingual | Claude Opus 4.585.7 | Margin← 2.6 | GLM-583.1 |
| Math | Claude Opus 4.557.5 | Margin← 1.2 | GLM-556.3 |
| Multimodal | Claude Opus 4.569.9 | MarginNo overlap | GLM-5Not measured |
Decisive benchmark drivers
The largest measured benchmark gaps in this matchup, with exact reported values.
More
- Source ↗
HLE
KnowledgeA 30.8%B 50.4%Winner: GLM-5Δ 19.6HLE: Claude Opus 4.5 scored 30.8%; GLM-5 scored 50.4%. GLM-5 wins this benchmark. - Source ↗
FrontierMath v2 (Tiers 1-3)
MathA 20.690%B 16.434%Winner: Claude Opus 4.5Δ 4.3FrontierMath v2 (Tiers 1-3): Claude Opus 4.5 scored 20.690%; GLM-5 scored 16.434%. Claude Opus 4.5 wins this benchmark. - Source ↗
SuperGPQA
KnowledgeA 70.6%B 66.8%Winner: Claude Opus 4.5Δ 3.8SuperGPQA: Claude Opus 4.5 scored 70.6%; GLM-5 scored 66.8%. Claude Opus 4.5 wins this benchmark. - Source ↗
MMLU-Pro
KnowledgeA 89.5%B 85.7%Winner: Claude Opus 4.5Δ 3.8MMLU-Pro: Claude Opus 4.5 scored 89.5%; GLM-5 scored 85.7%. Claude Opus 4.5 wins this benchmark. - Source ↗
LongBench v2
ReasoningA 64.4%B 60.8%Winner: Claude Opus 4.5Δ 3.6LongBench v2: Claude Opus 4.5 scored 64.4%; GLM-5 scored 60.8%. Claude Opus 4.5 wins this benchmark.
Operational comparison
Runtime and commercial metrics are compared only when both models have a complete sourced value.
| Metric | Claude Opus 4.5 | GLM-5 | Comparison |
|---|---|---|---|
| Input / output priceUSD per 1M tokens | Claude Opus 4.5$5 input / $25 output | GLM-5$1 input / $3.2 output | GLM-5 has the lower combined listed price. |
| Generation speedtokens per second | Claude Opus 4.546 tok/s | GLM-574 tok/s | GLM-5 has the higher measured throughput. |
| First-answer latencyseconds to first token | Claude Opus 4.51.01 s | GLM-51.64 s | Claude Opus 4.5 reaches the first token sooner. |
| Context windowmaximum listed tokens | Claude Opus 4.5200K | GLM-5200K | Listed context windows are equal. |
Benchmark Deep Dive
AgenticClaude Opus 4.5 wins17 benchmarks
| Benchmark | Claude Opus 4.5 | GLM-5 | Result |
|---|---|---|---|
| Terminal-Bench 2.0Source | 59.3% | 56.2% | Claude Opus 4.5 leads |
| OSWorld-VerifiedSource | 66.3% | — | Not comparable |
| OSWorldSource | 66.3% | — | Not comparable |
| Claw-EvalSource | 59.6% | 57.7% | Claude Opus 4.5 leads |
| QwenClawBenchSource | 52.3% | 54.1% | GLM-5 leads |
| τ³-bench resultsSource | 70.2% | 65.6% | Claude Opus 4.5 leads |
| VITA-BenchSource | 23.3% | — | Not comparable |
| DeepPlanningSource | 26.4% | 14.6% | Claude Opus 4.5 leads |
| ToolathlonSource | 43.5% | 38% | Claude Opus 4.5 leads |
| MCP AtlasSource | 42.3% | 31.1% | Claude Opus 4.5 leads |
| MCP-TasksSource | 71.8% | 60.8% | Claude Opus 4.5 leads |
| WideResearchSource | 76.4% | 69.8% | Claude Opus 4.5 leads |
| CyberGymSource | 50.6% | 43.2% | Claude Opus 4.5 leads |
| τ²-bench resultsSource | 86.3% | 98.2% | GLM-5 leads |
| Gert LabsSource | 64.23% | 50.99% | Claude Opus 4.5 leads |
| JobBenchSource | 32.3% | — | Not comparable |
| APEX-Agents-AASource | — | 14.5% | Not comparable |
CodingClaude Opus 4.5 wins9 benchmarks
| Benchmark | Claude Opus 4.5 | GLM-5 | Result |
|---|---|---|---|
| SWE-bench VerifiedSource | 80.9% | 77.8% | Claude Opus 4.5 leads |
| LiveCodeBench v6Source | 84.8% | — | Not comparable |
| SWE-bench ProSource | 57.1% | 55.1% | Claude Opus 4.5 leads |
| SWE MultilingualSource | 77.5% | 73.3% | Claude Opus 4.5 leads |
| NL2RepoSource | 43.2% | — | Not comparable |
| AA-SciCodeSource | 47.0% | 46.2% | Claude Opus 4.5 leads |
| SWE-bench Verified*Source | — | 72.8% | Not comparable |
| SWE-RebenchSource | — | 62.8% | Not comparable |
| React Native EvalsSource | — | 74.8% | Not comparable |
ReasoningClaude Opus 4.5 wins4 benchmarks
KnowledgeGLM-5 wins15 benchmarks
| Benchmark | Claude Opus 4.5 | GLM-5 | Result |
|---|---|---|---|
| GPQASource | 87% | 86% | Claude Opus 4.5 leads |
| SuperGPQASource | 70.6% | 66.8% | Claude Opus 4.5 leads |
| MMLU-ProSource | 89.5% | 85.7% | Claude Opus 4.5 leads |
| MMLU-ReduxSource | 96.6% | — | Not comparable |
| C-EvalSource | 92.2% | — | Not comparable |
| HLESource | 30.8% | 50.4% | GLM-5 leads |
| Artificial Analysis Intelligence IndexSource | 34.7% | 39.5% | GLM-5 leads |
| AA-GPQA DiamondSource | 81.0% | 82.0% | GLM-5 leads |
| AA-HLESource | 12.9% | 27.2% | GLM-5 leads |
| AA-Omniscience IndexSource | -3.9% | 2.0% | GLM-5 leads |
| AA-Omniscience AccuracySource | 40.7% | 26.9% | Claude Opus 4.5 leads |
| AA-Omniscience Hallucination RateSource | 75.4% | 34.0% | GLM-5 leads |
| AA MMLU-ProSource | 88.9% | — | Not comparable |
| GPQA-DSource | — | 86.0% | Not comparable |
| MMLU-Pro (Arcee)Source | — | 85.8% | Not comparable |
MathClaude Opus 4.5 wins8 benchmarks
| Benchmark | Claude Opus 4.5 | GLM-5 | Result |
|---|---|---|---|
| AIME26Source | 95.1% | 95.8% | GLM-5 leads |
| HMMT Feb 2025Source | 92.9% | 97.5% | GLM-5 leads |
| HMMT Nov 2025Source | 93.3% | 96.9% | GLM-5 leads |
| HMMT Feb 2026Source | 85.3% | 86.4% | GLM-5 leads |
| MMAnswerBenchSource | 84.0% | 82.5% | Claude Opus 4.5 leads |
| FrontierMath v2 (Tiers 1-3)Source | 20.690% | 16.434% | Claude Opus 4.5 leads |
| FrontierMath v2 (Tier 4)Source | 4.167% | 2.100% | Claude Opus 4.5 leads |
| AIME25 (Arcee)Source | — | 93.3% | Not comparable |
MultilingualClaude Opus 4.5 wins2 benchmarks
Multimodal8 benchmarks
| Benchmark | Claude Opus 4.5 | GLM-5 | Result |
|---|---|---|---|
| MMMU-ProSource | 70.6% | — | Not comparable |
| MathVisionSource | 74.3% | — | Not comparable |
| CharXivSource | 68.5% | — | Not comparable |
| VideoMMMUSource | 84.4% | — | Not comparable |
| ScreenSpot ProSource | 45.7% | — | Not comparable |
| V*Source | 67.0% | — | Not comparable |
| AA-MMMU-ProSource | 71.2% | — | Not comparable |
| Design Arena WebsiteSource | 1277 | 1278 | GLM-5 leads |
Frequently Asked Questions (8)
Which is better, Claude Opus 4.5 or GLM-5?
GLM-5 is ahead on BenchLM's BenchAlign leaderboard, 66.06 to 64.22. The biggest single separator in this matchup is HLE, where the scores are 30.8% and 50.4%.
Which is better for knowledge tasks, Claude Opus 4.5 or GLM-5?
GLM-5 has the edge for knowledge tasks in this comparison, averaging 66.4 versus 58.1. Inside this category, AA-Omniscience Hallucination Rate is the benchmark that creates the most daylight between them.
Which is better for coding, Claude Opus 4.5 or GLM-5?
Claude Opus 4.5 has the edge for coding in this comparison, averaging 71.7 versus 66.3. Inside this category, SWE Multilingual is the benchmark that creates the most daylight between them.
Which is better for math, Claude Opus 4.5 or GLM-5?
Claude Opus 4.5 has the edge for math in this comparison, averaging 57.5 versus 56.3. Inside this category, HMMT Feb 2025 is the benchmark that creates the most daylight between them.
Which is better for reasoning, Claude Opus 4.5 or GLM-5?
Claude Opus 4.5 has the edge for reasoning in this comparison, averaging 64.4 versus 60.8. Inside this category, AI-Needle is the benchmark that creates the most daylight between them.
Which is better for agentic tasks, Claude Opus 4.5 or GLM-5?
Claude Opus 4.5 has the edge for agentic tasks in this comparison, averaging 62.6 versus 56.2. Inside this category, Gert Labs is the benchmark that creates the most daylight between them.
Which is better for instruction following, Claude Opus 4.5 or GLM-5?
GLM-5 has the edge for instruction following in this comparison, averaging 92.6 versus 69.5. Inside this category, AA-IFBench is the benchmark that creates the most daylight between them.
Which is better for multilingual tasks, Claude Opus 4.5 or GLM-5?
Claude Opus 4.5 has the edge for multilingual tasks in this comparison, averaging 85.7 versus 83.1. Inside this category, MMLU-ProX is the benchmark that creates the most daylight between them.
Related Comparisons
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.