Model comparison
GLM-4.7 vs MiniMax M2.7
Head-to-head evidence from 20 shared benchmark results across 6 categories. Overall scores shown here use BenchLM's provisional ranking lane.
Verified leaderboard positions: GLM-4.7 #32; MiniMax M2.7 unranked
Evidence parity. GLM-4.7 and MiniMax M2.7 share 20 comparable benchmark results. 2 of 8 categories are comparable. 11 results are unique to GLM-4.7; 17 to MiniMax M2.7.
Updated July 14, 2026- Shared results
- 20
- GLM-4.7 only
- 11
- MiniMax M2.7 only
- 17
- Comparable categories
- 2 / 8
Pick GLM-4.7 if you want the stronger benchmark profile. MiniMax M2.7 only becomes the better choice if agentic is the priority or you would rather avoid the extra latency and token burn of a reasoning model.
Confidence note. This is a partial-evidence comparison with 20 shared benchmark results across 6 evidence categories; 2 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.
Why this result
GLM-4.7 is clearly ahead on the provisional aggregate, 62 to 55. The gap is large enough that you do not need to squint at the spreadsheet to see the difference.
GLM-4.7's sharpest advantage is in coding, where it averages 73.8 against 54.4. The single biggest benchmark swing on the page is Terminal-Bench 2.0, 41% to 57%. MiniMax M2.7 does hit back in agentic, so the answer changes if that is the part of the workload you care about most.
MiniMax M2.7 is also the more expensive model on tokens at $0.30 input / $1.20 output per 1M tokens, versus $0.00 input / $0.00 output per 1M tokens for GLM-4.7. That is roughly Infinityx on output cost alone. GLM-4.7 is the reasoning model in the pair, while MiniMax M2.7 is not. That usually helps on harder chain-of-thought-heavy tests, but it can also mean more latency and more token spend in real use.
Category breakdown
Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.
| Category | GLM-4.7 | Δ | MiniMax M2.7 |
|---|---|---|---|
| Coding | GLM-4.773.8 | Margin← 19.4 | MiniMax M2.754.4 |
| Agentic | GLM-4.745.7 | Margin→ 11.3 | MiniMax M2.757.0 |
| Knowledge | GLM-4.752.1 | MarginNo overlap | MiniMax M2.7Not measured |
| Math | GLM-4.71.8 | MarginNo overlap | MiniMax M2.7Not measured |
Decisive benchmark drivers
The largest measured benchmark gaps in this matchup, with exact reported values.
More
- Source ↗
Terminal-Bench 2.0
AgenticA 41%B 57%Winner: MiniMax M2.7Δ 16Terminal-Bench 2.0: GLM-4.7 scored 41%; MiniMax M2.7 scored 57%. MiniMax M2.7 wins this benchmark. - Source ↗
SWE-Rebench
CodingA 58.7%B 51.9%Winner: GLM-4.7Δ 6.8SWE-Rebench: GLM-4.7 scored 58.7%; MiniMax M2.7 scored 51.9%. GLM-4.7 wins this benchmark.
Operational comparison
Runtime and commercial metrics are compared only when both models have a complete sourced value.
| Metric | GLM-4.7 | MiniMax M2.7 | Comparison |
|---|---|---|---|
| Input / output priceUSD per 1M tokens | GLM-4.7$0 input / $0 output | MiniMax M2.7$0.3 input / $1.2 output | GLM-4.7 has the lower combined listed price. |
| Generation speedtokens per second | GLM-4.782 tok/s | MiniMax M2.745 tok/s | GLM-4.7 has the higher measured throughput. |
| First-answer latencyseconds to first token | GLM-4.71.10 s | MiniMax M2.72.53 s | GLM-4.7 reaches the first token sooner. |
| Context windowmaximum listed tokens | GLM-4.7200K | MiniMax M2.7200K | Listed context windows are equal. |
Benchmark Deep Dive
AgenticMiniMax M2.7 wins13 benchmarks
| Benchmark | GLM-4.7 | MiniMax M2.7 | Result |
|---|---|---|---|
| Terminal-Bench 2.0Source | 41% | 57% | MiniMax M2.7 leads |
| BrowseCompSource | 52% | — | Not comparable |
| VITA-BenchSource | 15.5% | — | Not comparable |
| AA Agentic IndexSource | 25.4% | 25.6% | MiniMax M2.7 leads |
| Tau2-TelecomSource | 95.9% | 84.8% | GLM-4.7 leads |
| Gert LabsSource | 39.95% | 40.40% | MiniMax M2.7 leads |
| GDPval-AASource | 33.3% | 32.9% | GLM-4.7 leads |
| GDPval-AASource | 1165 | 1158 | GLM-4.7 leads |
| ToolathlonSource | — | 46.3% | Not comparable |
| MLE-Bench LiteSource | — | 66.6% | Not comparable |
| MM-ClawBenchSource | — | 62.7% | Not comparable |
| Claw-EvalSource | — | 48.7% | Not comparable |
| APEX-Agents-AASource | — | 10.6% | Not comparable |
CodingGLM-4.7 wins15 benchmarks
| Benchmark | GLM-4.7 | MiniMax M2.7 | Result |
|---|---|---|---|
| SWE-bench VerifiedSource | 73.8% | — | Not comparable |
| LiveCodeBenchSource | 84.9% | — | Not comparable |
| SWE-RebenchSource | 58.7% | 51.9% | GLM-4.7 leads |
| AA Coding IndexSource | 45.3% | 52.6% | MiniMax M2.7 leads |
| Terminal-Bench HardSource | 31.8% | 39.4% | MiniMax M2.7 leads |
| AA-SciCodeSource | 45.1% | 47.0% | MiniMax M2.7 leads |
| AA LiveCodeBenchSource | 89.4% | — | Not comparable |
| SWE-bench Verified*Source | — | 75.4% | Not comparable |
| SWE-bench ProSource | — | 56.2% | Not comparable |
| SWE MultilingualSource | — | 76.5% | Not comparable |
| Multi-SWE BenchSource | — | 52.7% | Not comparable |
| VIBE-ProSource | — | 55.6% | Not comparable |
| NL2RepoSource | — | 39.8% | Not comparable |
| Vibe Code BenchSource | — | 27.04% | Not comparable |
| React Native EvalsSource | — | 71.4% | Not comparable |
Reasoning2 benchmarks
Knowledge11 benchmarks
| Benchmark | GLM-4.7 | MiniMax M2.7 | Result |
|---|---|---|---|
| GPQASource | 85.7% | — | Not comparable |
| MMLU-ProSource | 84.3% | — | Not comparable |
| HLESource | 24.8% | — | Not comparable |
| Artificial Analysis Intelligence IndexSource | 33.7% | 38.1% | MiniMax M2.7 leads |
| AA-GPQA DiamondSource | 85.9% | 87.4% | MiniMax M2.7 leads |
| AA-HLESource | 25.1% | 28.1% | MiniMax M2.7 leads |
| AA-Omniscience IndexSource | -34.6% | 0.7% | MiniMax M2.7 leads |
| AA-Omniscience AccuracySource | 29.3% | 26.1% | GLM-4.7 leads |
| AA-Omniscience Hallucination RateSource | 90.3% | 34.4% | MiniMax M2.7 leads |
| GPQA-DSource | — | 87.0% | Not comparable |
| MMLU-Pro (Arcee)Source | — | 80.8% | Not comparable |
Math4 benchmarks
Multimodal2 benchmarks
Inst. Following1 benchmarks
| Benchmark | GLM-4.7 | MiniMax M2.7 | Result |
|---|---|---|---|
| AA-IFBenchSource | 67.9% | 75.7% | MiniMax M2.7 leads |
Frequently Asked Questions (3)
Which is better, GLM-4.7 or MiniMax M2.7?
GLM-4.7 is ahead on BenchLM's provisional leaderboard, 62 to 55. The biggest single separator in this matchup is Terminal-Bench 2.0, where the scores are 41% and 57%.
Which is better for coding, GLM-4.7 or MiniMax M2.7?
GLM-4.7 has the edge for coding in this comparison, averaging 73.8 versus 54.4. Inside this category, Terminal-Bench Hard is the benchmark that creates the most daylight between them.
Which is better for agentic tasks, GLM-4.7 or MiniMax M2.7?
MiniMax M2.7 has the edge for agentic tasks in this comparison, averaging 57 versus 45.7. Inside this category, Terminal-Bench 2.0 is the benchmark that creates the most daylight between them.
Related Comparisons
Explore More
The AI models change fast. We track them for you.
A weekly brief for engineers and researchers covering new models, ranking shifts, and pricing changes.
Free. No spam. Unsubscribe anytime.