Skip to main content

Model comparison

Mistral Medium 3.5 128B vs Nemotron 3 Ultra

Data verified

Head-to-head evidence from 23 shared benchmark results across 5 categories. Overall scores shown here use BenchLM's provisional ranking lane.

Margin
1.0pts
winning →
63/100
1 category wins0 category wins

Verified leaderboard positions: Mistral Medium 3.5 128B unranked; Nemotron 3 Ultra #21

Evidence parity. Mistral Medium 3.5 128B and Nemotron 3 Ultra share 23 comparable benchmark results. 1 of 8 categories are comparable. 2 results are unique to Mistral Medium 3.5 128B; 17 to Nemotron 3 Ultra.

Updated July 14, 2026
Shared results
23
Mistral Medium 3.5 128B only
2
Nemotron 3 Ultra only
17
Comparable categories
1 / 8

Pick Nemotron 3 Ultra if you want the stronger benchmark profile. Mistral Medium 3.5 128B only becomes the better choice if coding is the priority.

Confidence note. This is a partial-evidence comparison with 23 shared benchmark results across 5 evidence categories; 1 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.

Why this result

Nemotron 3 Ultra finishes one point ahead on BenchLM's provisional leaderboard, 63 to 62. That is enough to call, but not enough to treat as a blowout. This matchup comes down to a few meaningful edges rather than one model dominating the board.

Mistral Medium 3.5 128B is also the more expensive model on tokens at $1.50 input / $7.50 output per 1M tokens, versus $0.00 input / $0.00 output per 1M tokens for Nemotron 3 Ultra. That is roughly Infinityx on output cost alone. Nemotron 3 Ultra gives you the larger context window at 1M, compared with 256K for Mistral Medium 3.5 128B.

Category breakdown

Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.

Category scores and score margins for Mistral Medium 3.5 128B and Nemotron 3 Ultra
CategoryMistral Medium 3.5 128BΔNemotron 3 Ultra
CodingMistral Medium 3.5 128B77.6Margin 2.1Nemotron 3 Ultra75.5
AgenticMistral Medium 3.5 128BNot measuredMarginNo overlapNemotron 3 Ultra51.3
ReasoningMistral Medium 3.5 128BNot measuredMarginNo overlapNemotron 3 Ultra61.9
KnowledgeMistral Medium 3.5 128BNot measuredMarginNo overlapNemotron 3 Ultra54.2
MultilingualMistral Medium 3.5 128BNot measuredMarginNo overlapNemotron 3 Ultra83.0
Inst. FollowingMistral Medium 3.5 128BNot measuredMarginNo overlapNemotron 3 Ultra81.7

Decisive benchmark drivers

The largest measured benchmark gaps in this matchup, with exact reported values.

More
A · Mistral Medium 3.5 128BB · Nemotron 3 Ultra
  1. SWE-bench Verified

    Coding
    Source ↗
    A 77.6%B 71.9%
    Winner: Mistral Medium 3.5 128BΔ 5.7
    SWE-bench Verified: Mistral Medium 3.5 128B scored 77.6%; Nemotron 3 Ultra scored 71.9%. Mistral Medium 3.5 128B wins this benchmark.

Operational comparison

Runtime and commercial metrics are compared only when both models have a complete sourced value.

MetricMistral Medium 3.5 128BNemotron 3 UltraComparison
Input / output priceUSD per 1M tokensMistral Medium 3.5 128B$1.5 input / $7.5 outputNemotron 3 Ultra$0 input / $0 outputNemotron 3 Ultra has the lower combined listed price.
Generation speedtokens per secondMistral Medium 3.5 128BNot availableNemotron 3 UltraNot availableA complete speed comparison is not available.
First-answer latencyseconds to first tokenMistral Medium 3.5 128BNot availableNemotron 3 UltraNot availableA complete latency comparison is not available.
Context windowmaximum listed tokensMistral Medium 3.5 128B256KNemotron 3 Ultra1MNemotron 3 Ultra lists the larger context window.

Benchmark Deep Dive

Agentic
BenchmarkMistral Medium 3.5 128BNemotron 3 UltraResult
TAU3-BenchSource 91.4%70.9%Mistral Medium 3.5 128B leads
AA Agentic IndexSource 19.0%27.4%Nemotron 3 Ultra leads
Tau2-TelecomSource 94.2%83.3%Mistral Medium 3.5 128B leads
GDPval-AASource 21.4%33.2%Nemotron 3 Ultra leads
GDPval-AASource 9291164Nemotron 3 Ultra leads
Gert LabsSource 39.10%Not comparable
AA BriefcaseSource 506870Nemotron 3 Ultra leads
AA EnterpriseOps-GymSource 33.7%28.9%Mistral Medium 3.5 128B leads
AA Harvey LABSource 0.8%3.3%Nemotron 3 Ultra leads
AA Tau3 BankingSource 14.4%13.8%Mistral Medium 3.5 128B leads
Terminal-Bench 2.0Source 56.4%Not comparable
PinchBenchSource 90.0%Not comparable
BrowseCompSource 44.4%Not comparable
HLE w/ toolsSource 37.4%Not comparable
CodingMistral Medium 3.5 128B wins
BenchmarkMistral Medium 3.5 128BNemotron 3 UltraResult
SWE-bench VerifiedSource 77.6%71.9%Mistral Medium 3.5 128B leads
AA Coding IndexSource 46.9%49.3%Nemotron 3 Ultra leads
Terminal-Bench HardSource 33.3%36.4%Nemotron 3 Ultra leads
AA-SciCodeSource 39.6%39.9%Nemotron 3 Ultra leads
SWE MultilingualSource 67.7%Not comparable
LiveCodeBenchSource 89%Not comparable
SciCodeSource 44.6%Not comparable
Terminal-Bench 2.0Source 56.4%Not comparable
Reasoning
BenchmarkMistral Medium 3.5 128BNemotron 3 UltraResult
AA-LCRSource 61.0%67.0%Nemotron 3 Ultra leads
CritPtSource 0.0%3.1%Nemotron 3 Ultra leads
LongBench v2Source 61.9%Not comparable
Knowledge
BenchmarkMistral Medium 3.5 128BNemotron 3 UltraResult
Artificial Analysis Intelligence IndexSource 29.9%37.8%Nemotron 3 Ultra leads
AA-GPQA DiamondSource 74.8%86.7%Nemotron 3 Ultra leads
AA-HLESource 12.8%26.6%Nemotron 3 Ultra leads
AA-Omniscience IndexSource -36.3%-0.8%Nemotron 3 Ultra leads
AA-Omniscience AccuracySource 25.1%21.6%Mistral Medium 3.5 128B leads
AA-Omniscience Hallucination RateSource 82.0%28.5%Nemotron 3 Ultra leads
AA Openness IndexSource 33.3%83.3%Nemotron 3 Ultra leads
GPQASource 87%Not comparable
GPQA-DSource 87.0%Not comparable
HLESource 26.7%Not comparable
HLE w/o toolsSource 26.7%Not comparable
MMLU-ProSource 86.8%Not comparable
Multilingual
BenchmarkMistral Medium 3.5 128BNemotron 3 UltraResult
MMLU-ProXSource 83%Not comparable
Multimodal
BenchmarkMistral Medium 3.5 128BNemotron 3 UltraResult
AA-MMMU-ProSource 64.9%Not comparable
Design Arena WebsiteSource 1132Not comparable
Inst. Following
BenchmarkMistral Medium 3.5 128BNemotron 3 UltraResult
AA-IFBenchSource 68.8%81.4%Nemotron 3 Ultra leads
IFBenchSource 81.7%Not comparable
Frequently Asked Questions (2)

Which is better, Mistral Medium 3.5 128B or Nemotron 3 Ultra?

Nemotron 3 Ultra is ahead on BenchLM's provisional leaderboard, 63 to 62. The biggest single separator in this matchup is SWE-bench Verified, where the scores are 77.6% and 71.9%.

Which is better for coding, Mistral Medium 3.5 128B or Nemotron 3 Ultra?

Mistral Medium 3.5 128B has the edge for coding in this comparison, averaging 77.6 versus 75.5. Inside this category, SWE-bench Verified is the benchmark that creates the most daylight between them.

Related Comparisons

Last updated: July 14, 2026

The AI models change fast. We track them for you.

A weekly brief for engineers and researchers covering new models, ranking shifts, and pricing changes.

Free. No spam. Unsubscribe anytime.