# Kimi K3 vs Mistral Medium 3.5 128B

> Side-by-side benchmark comparison for Kimi K3 and Mistral Medium 3.5 128B across agentic, coding, multimodal, knowledge, reasoning, multilingual, and math tasks.

- Canonical page: https://benchlm.ai/compare/kimi-k3-vs-mistral-medium-3-5-128b
- Last updated: October 2, 2026

- Shared sourced benchmarks: 4
- HTML indexing: indexable
- Ranking lane: BenchAlign v5.8

## Quick Verdict

Kimi K3 leads by point estimate. Conditional score ranges do not establish rank confidence. Choose using the category evidence and your own workload tests.

## Summary

- Kimi K3 has the higher overall point estimate 72.14 to 36.17. Conditional score ranges do not establish rank confidence.
- The clearest category separation is in agentic, where the averages are 68.1 for Kimi K3 and 19.3 for Mistral Medium 3.5 128B.
- The biggest single benchmark swing is GPQA Diamond (Vals) in Knowledge, with scores of 92.9% and 34.8%.
- Mistral Medium 3.5 128B is the cheaper option on output tokens, which matters if you expect large responses or heavy interactive use.
- Kimi K3 also has the larger context window at 1.05M.

## Model Snapshot

| Property | Kimi K3 | Mistral Medium 3.5 128B |
|----------|----------|----------|
| Creator | Moonshot AI | Mistral |
| Type | Pending | Open Weight |
| Reasoning | Reasoning | Reasoning |
| Context | 1.05M | 256K |
| Overall Score | 72.14 | 36.17 |
| Benchmarks Covered | 48 | 7 |
| Pricing (input/output) | $3.00 / $15.00 | $1.50 / $7.50 |

## Category Breakdown

### Agentic

- Winner: Kimi K3
- Kimi K3 public-lane score: 68.1 (Supported · #8/119)
- Mistral Medium 3.5 128B public-lane score: 19.3 (Supported · #99/119)

| Benchmark | Kimi K3 | Mistral Medium 3.5 128B | Winner |
|-----------|-----------|-----------|--------|
| Terminal-Bench 2.1 | 88.3% | Coming soon | Coming soon |
| BrowseComp | 91.2% | Coming soon | Coming soon |
| DeepSearchQA | 95.0% | Coming soon | Coming soon |
| Toolathlon-Verified | 73.2% | Coming soon | Coming soon |
| MCP Atlas | 84.2% | Coming soon | Coming soon |
| AutomationBench | 30.8% | Coming soon | Coming soon |
| JobBench | 52.9% | Coming soon | Coming soon |
| APEX-Agents | 37.6% | Coming soon | Coming soon |
| SpreadsheetBench 2 | 34.8% | Coming soon | Coming soon |
| DECK-Bench | 73.5% | Coming soon | Coming soon |
| Terminal-Bench 2.1 (Vals) | 80.9% | 39.0% | Kimi K3 |
| ApprenticeBench | 18% | Coming soon | Coming soon |
| τ³-bench results | Coming soon | 91.4% | Coming soon |
| Gert Labs | Coming soon | 39.10% | Coming soon |

### Coding

- Winner: Directional only
- Kimi K3 public-lane score: 61.4 (Supported · #18/144)
- Mistral Medium 3.5 128B public-lane score: 25.3 (Estimated · #105/144)

| Benchmark | Kimi K3 | Mistral Medium 3.5 128B | Winner |
|-----------|-----------|-----------|--------|
| DeepSWE | 67.5% | Coming soon | Coming soon |
| CursorBench 3.2 | 60.8% | Coming soon | Coming soon |
| FrontierSWE | 81.2% | Coming soon | Coming soon |
| ProgramBench | 77.8% | Coming soon | Coming soon |
| Kimi Code Bench v2 | 72.9% | Coming soon | Coming soon |
| sweMarathon | 42% | Coming soon | Coming soon |
| PostTrain Bench | 36.6% | Coming soon | Coming soon |
| MLS-Bench Lite | 48.3% | Coming soon | Coming soon |
| VulcanBench v3 | 73.7% | Coming soon | Coming soon |
| OpenHarmony Bench | 57.3% | Coming soon | Coming soon |
| FrontierSWE v2 | 25.9% | Coming soon | Coming soon |
| LiveCodeBench (Vals) | 87.2% | Coming soon | Coming soon |
| SWE-bench (Vals) | 93.4% | 66.4% | Kimi K3 |
| PostTrainBench v1.1 | 32.0% | Coming soon | Coming soon |
| SWE-bench Verified | Coming soon | 77.6% | Coming soon |

### Reasoning

- Winner: Not comparable
- Kimi K3 public-lane score: 65.8 (#18/27)
- Mistral Medium 3.5 128B public-lane score: 69.9 (Unranked · 2 rankable rows)

| Benchmark | Kimi K3 | Mistral Medium 3.5 128B | Winner |
|-----------|-----------|-----------|--------|
| ARC-AGI-1 | 94.50% | Coming soon | Coming soon |
| ARC-AGI-2 | 60.4% | Coming soon | Coming soon |

### Multimodal & Grounded

- Winner: Not comparable
- Kimi K3 public-lane score: 89.4 (#1/49)
- Mistral Medium 3.5 128B public-lane score: 56.7 (Unranked · 1 rankable row)

| Benchmark | Kimi K3 | Mistral Medium 3.5 128B | Winner |
|-----------|-----------|-----------|--------|
| OfficeQA Pro | 63.3% | Coming soon | Coming soon |
| MMMU-Pro | 81.6% | Coming soon | Coming soon |
| MMMU-Pro w/ Python | 83.4% | Coming soon | Coming soon |
| CharXiv w/o tools | 84.8% | Coming soon | Coming soon |
| CharXiv | 91.3% | Coming soon | Coming soon |
| MathVision | 94.3% | Coming soon | Coming soon |
| MathVision w/ Python | 97.8% | Coming soon | Coming soon |
| BabyVision w/ Python | 85.7% | Coming soon | Coming soon |
| ZeroBench | 23.0% | Coming soon | Coming soon |
| ZeroBench w/ Python | 41.0% | Coming soon | Coming soon |
| WorldVQA ForceAnswer | 51.0% | Coming soon | Coming soon |
| OmniDocBench | 91.1% | Coming soon | Coming soon |
| PerceptionBench | 58.5% | Coming soon | Coming soon |

### Knowledge

- Winner: Kimi K3
- Kimi K3 public-lane score: 67.9 (Supported · #18/171)
- Mistral Medium 3.5 128B public-lane score: 32.9 (Supported · #124/171)

| Benchmark | Kimi K3 | Mistral Medium 3.5 128B | Winner |
|-----------|-----------|-----------|--------|
| GPQA | 93.5% | Coming soon | Coming soon |
| GPQA-D | 93.5% | Coming soon | Coming soon |
| HLE | 56% | Coming soon | Coming soon |
| HLE w/o tools | 43.5% | Coming soon | Coming soon |
| GPQA Diamond (Vals) | 92.9% | 34.8% | Kimi K3 |
| MMLU-Pro (Vals) | 88.0% | 75.3% | Kimi K3 |

### Instruction Following

- Winner: Not comparable
- Kimi K3 public-lane score: Coming soon
- Mistral Medium 3.5 128B public-lane score: 82.6 (#46/125)

| Benchmark | Kimi K3 | Mistral Medium 3.5 128B | Winner |
|-----------|-----------|-----------|--------|
| Gray Swan IPI (15 attempts) | 52.7% | Coming soon | Coming soon |

## FAQ

### Which is better overall, Kimi K3 or Mistral Medium 3.5 128B?

Kimi K3 has the higher overall point estimate. Conditional score ranges do not establish rank confidence.

### Where is the biggest gap between Kimi K3 and Mistral Medium 3.5 128B?

The widest category gap is in agentic, where the averages are 68.1 for Kimi K3 and 19.3 for Mistral Medium 3.5 128B.

### How many benchmarks does Kimi K3 cover on BenchLM?

Kimi K3 currently has 48 sourced benchmark scores on BenchLM.

### How many benchmarks does Mistral Medium 3.5 128B cover on BenchLM?

Mistral Medium 3.5 128B currently has 7 sourced benchmark scores on BenchLM.

## Related Comparisons

- [Kimi K3 vs GPT-6 Astra](/compare/gpt-6-astra-vs-kimi-k3)
- [Mistral Medium 3.5 128B vs GPT-6 Astra](/compare/gpt-6-astra-vs-mistral-medium-3-5-128b)
- [Kimi K3 vs Claude Opus 5.5](/compare/claude-opus-5-5-vs-kimi-k3)
- [Mistral Medium 3.5 128B vs Claude Opus 5.5](/compare/claude-opus-5-5-vs-mistral-medium-3-5-128b)

## Explore More

- [Kimi K3 profile](/models/kimi-k3)
- [Mistral Medium 3.5 128B profile](/models/mistral-medium-3-5-128b)
- [Compare Pricing](/llm-pricing)
- [Alternative Finder](/tools/alternative-finder)
- [LLM Selector](/tools/llm-selector)
- [Overall Rankings](/best/overall)
