Skip to main content

Model comparison

GLM-4.7 vs MiniMax M2.7

Data verified

Head-to-head evidence from 20 shared benchmark results across 6 categories. Overall scores shown here use BenchLM's provisional ranking lane.

62/100
Margin
7.0pts
← winning
55/100
1 category wins1 category wins

Verified leaderboard positions: GLM-4.7 #32; MiniMax M2.7 unranked

Evidence parity. GLM-4.7 and MiniMax M2.7 share 20 comparable benchmark results. 2 of 8 categories are comparable. 11 results are unique to GLM-4.7; 17 to MiniMax M2.7.

Updated July 14, 2026
Shared results
20
GLM-4.7 only
11
MiniMax M2.7 only
17
Comparable categories
2 / 8

Pick GLM-4.7 if you want the stronger benchmark profile. MiniMax M2.7 only becomes the better choice if agentic is the priority or you would rather avoid the extra latency and token burn of a reasoning model.

Confidence note. This is a partial-evidence comparison with 20 shared benchmark results across 6 evidence categories; 2 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.

Why this result

GLM-4.7 is clearly ahead on the provisional aggregate, 62 to 55. The gap is large enough that you do not need to squint at the spreadsheet to see the difference.

GLM-4.7's sharpest advantage is in coding, where it averages 73.8 against 54.4. The single biggest benchmark swing on the page is Terminal-Bench 2.0, 41% to 57%. MiniMax M2.7 does hit back in agentic, so the answer changes if that is the part of the workload you care about most.

MiniMax M2.7 is also the more expensive model on tokens at $0.30 input / $1.20 output per 1M tokens, versus $0.00 input / $0.00 output per 1M tokens for GLM-4.7. That is roughly Infinityx on output cost alone. GLM-4.7 is the reasoning model in the pair, while MiniMax M2.7 is not. That usually helps on harder chain-of-thought-heavy tests, but it can also mean more latency and more token spend in real use.

Category breakdown

Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.

Category scores and score margins for GLM-4.7 and MiniMax M2.7
CategoryGLM-4.7ΔMiniMax M2.7
CodingGLM-4.773.8Margin 19.4MiniMax M2.754.4
AgenticGLM-4.745.7Margin 11.3MiniMax M2.757.0
KnowledgeGLM-4.752.1MarginNo overlapMiniMax M2.7Not measured
MathGLM-4.71.8MarginNo overlapMiniMax M2.7Not measured

Decisive benchmark drivers

The largest measured benchmark gaps in this matchup, with exact reported values.

More
A · GLM-4.7B · MiniMax M2.7
  1. Terminal-Bench 2.0

    Agentic
    Source ↗
    A 41%B 57%
    Winner: MiniMax M2.7Δ 16
    Terminal-Bench 2.0: GLM-4.7 scored 41%; MiniMax M2.7 scored 57%. MiniMax M2.7 wins this benchmark.
  2. SWE-Rebench

    Coding
    Source ↗
    A 58.7%B 51.9%
    Winner: GLM-4.7Δ 6.8
    SWE-Rebench: GLM-4.7 scored 58.7%; MiniMax M2.7 scored 51.9%. GLM-4.7 wins this benchmark.

Operational comparison

Runtime and commercial metrics are compared only when both models have a complete sourced value.

MetricGLM-4.7MiniMax M2.7Comparison
Input / output priceUSD per 1M tokensGLM-4.7$0 input / $0 outputMiniMax M2.7$0.3 input / $1.2 outputGLM-4.7 has the lower combined listed price.
Generation speedtokens per secondGLM-4.782 tok/sMiniMax M2.745 tok/sGLM-4.7 has the higher measured throughput.
First-answer latencyseconds to first tokenGLM-4.71.10 sMiniMax M2.72.53 sGLM-4.7 reaches the first token sooner.
Context windowmaximum listed tokensGLM-4.7200KMiniMax M2.7200KListed context windows are equal.

Benchmark Deep Dive

AgenticMiniMax M2.7 wins
BenchmarkGLM-4.7MiniMax M2.7Result
Terminal-Bench 2.0Source 41%57%MiniMax M2.7 leads
BrowseCompSource 52%Not comparable
VITA-BenchSource 15.5%Not comparable
AA Agentic IndexSource 25.4%25.6%MiniMax M2.7 leads
Tau2-TelecomSource 95.9%84.8%GLM-4.7 leads
Gert LabsSource 39.95%40.40%MiniMax M2.7 leads
GDPval-AASource 33.3%32.9%GLM-4.7 leads
GDPval-AASource 11651158GLM-4.7 leads
ToolathlonSource 46.3%Not comparable
MLE-Bench LiteSource 66.6%Not comparable
MM-ClawBenchSource 62.7%Not comparable
Claw-EvalSource 48.7%Not comparable
APEX-Agents-AASource 10.6%Not comparable
CodingGLM-4.7 wins
BenchmarkGLM-4.7MiniMax M2.7Result
SWE-bench VerifiedSource 73.8%Not comparable
LiveCodeBenchSource 84.9%Not comparable
SWE-RebenchSource 58.7%51.9%GLM-4.7 leads
AA Coding IndexSource 45.3%52.6%MiniMax M2.7 leads
Terminal-Bench HardSource 31.8%39.4%MiniMax M2.7 leads
AA-SciCodeSource 45.1%47.0%MiniMax M2.7 leads
AA LiveCodeBenchSource 89.4%Not comparable
SWE-bench Verified*Source 75.4%Not comparable
SWE-bench ProSource 56.2%Not comparable
SWE MultilingualSource 76.5%Not comparable
Multi-SWE BenchSource 52.7%Not comparable
VIBE-ProSource 55.6%Not comparable
NL2RepoSource 39.8%Not comparable
Vibe Code BenchSource 27.04%Not comparable
React Native EvalsSource 71.4%Not comparable
Reasoning
BenchmarkGLM-4.7MiniMax M2.7Result
AA-LCRSource 64.0%68.7%MiniMax M2.7 leads
CritPtSource 1.7%0.6%GLM-4.7 leads
Knowledge
BenchmarkGLM-4.7MiniMax M2.7Result
GPQASource 85.7%Not comparable
MMLU-ProSource 84.3%Not comparable
HLESource 24.8%Not comparable
Artificial Analysis Intelligence IndexSource 33.7%38.1%MiniMax M2.7 leads
AA-GPQA DiamondSource 85.9%87.4%MiniMax M2.7 leads
AA-HLESource 25.1%28.1%MiniMax M2.7 leads
AA-Omniscience IndexSource -34.6%0.7%MiniMax M2.7 leads
AA-Omniscience AccuracySource 29.3%26.1%GLM-4.7 leads
AA-Omniscience Hallucination RateSource 90.3%34.4%MiniMax M2.7 leads
GPQA-DSource 87.0%Not comparable
MMLU-Pro (Arcee)Source 80.8%Not comparable
Math
BenchmarkGLM-4.7MiniMax M2.7Result
AIME 2025Source 95.7%Not comparable
FrontierMath v2 (Tiers 1-3)Source 2.439%Not comparable
FrontierMath v2 (Tier 4)Source 0.000%Not comparable
AIME25 (Arcee)Source 80.0%Not comparable
Multimodal
BenchmarkGLM-4.7MiniMax M2.7Result
Design Arena WebsiteSource 12601279MiniMax M2.7 leads
GDPval-AASource 1495Not comparable
Inst. Following
BenchmarkGLM-4.7MiniMax M2.7Result
AA-IFBenchSource 67.9%75.7%MiniMax M2.7 leads
Frequently Asked Questions (3)

Which is better, GLM-4.7 or MiniMax M2.7?

GLM-4.7 is ahead on BenchLM's provisional leaderboard, 62 to 55. The biggest single separator in this matchup is Terminal-Bench 2.0, where the scores are 41% and 57%.

Which is better for coding, GLM-4.7 or MiniMax M2.7?

GLM-4.7 has the edge for coding in this comparison, averaging 73.8 versus 54.4. Inside this category, Terminal-Bench Hard is the benchmark that creates the most daylight between them.

Which is better for agentic tasks, GLM-4.7 or MiniMax M2.7?

MiniMax M2.7 has the edge for agentic tasks in this comparison, averaging 57 versus 45.7. Inside this category, Terminal-Bench 2.0 is the benchmark that creates the most daylight between them.

Related Comparisons

Last updated: July 14, 2026

The AI models change fast. We track them for you.

A weekly brief for engineers and researchers covering new models, ranking shifts, and pricing changes.

Free. No spam. Unsubscribe anytime.