# Claude Opus 5.5 vs Kimi K3

> Side-by-side benchmark comparison for Claude Opus 5.5 and Kimi K3 across agentic, coding, multimodal, knowledge, reasoning, multilingual, and math tasks.

- Canonical page: https://benchlm.ai/compare/claude-opus-5-5-vs-kimi-k3
- Last updated: September 22, 2026

- Shared sourced benchmarks: 7
- HTML indexing: indexable
- Ranking lane: BenchAlign v5

## Quick Verdict

Pick Claude Opus 5.5 if you want the stronger benchmark profile. Kimi K3 only makes more sense when its price, context window, or workload-specific category wins matter more than the overall score.

## Summary

- Claude Opus 5.5 leads overall 81.03 to 73.99.
- The clearest category separation is in coding, where the averages are 80.1 for Claude Opus 5.5 and 66.5 for Kimi K3.
- The biggest single benchmark swing is FrontierSWE v2 in Coding, with scores of 62.3% and 25.9%.
- Kimi K3 is the cheaper option on output tokens, which matters if you expect large responses or heavy interactive use.
- Kimi K3 also has the larger context window at 1.05M.

## Model Snapshot

| Property | Claude Opus 5.5 | Kimi K3 |
|----------|----------|----------|
| Creator | Anthropic | Moonshot AI |
| Type | Proprietary | Pending |
| Reasoning | Reasoning | Reasoning |
| Context | 1M | 1.05M |
| Overall Score | 81.03 | 73.99 |
| Benchmarks Covered | 47 | 44 |
| Pricing (input/output) | $4.00 / $20.00 | $3.00 / $15.00 |

## Category Breakdown

### Agentic

- Winner: Claude Opus 5.5
- Claude Opus 5.5 public-lane score: 82.9 (Supported · #1/157)
- Kimi K3 public-lane score: 72.2 (Supported · #5/157)

| Benchmark | Claude Opus 5.5 | Kimi K3 | Winner |
|-----------|-----------|-----------|--------|
| Terminal-Bench 4.0 | 66.40% | Coming soon | Coming soon |
| Terminal-Bench-Science 0.1 | 58.7% | Coming soon | Coming soon |
| AutomationBench | 40.0% | 30.8% | Claude Opus 5.5 |
| HLE w/ tools | 67.7% | Coming soon | Coming soon |
| OSWorld 2.0 | 48.7% | Coming soon | Coming soon |
| LAB all-pass (Harvey held-out) | 8.3% | Coming soon | Coming soon |
| LAB criterion-pass (Harvey held-out) | 91.2% | Coming soon | Coming soon |
| Toolathlon-Verified | 77.8% | 73.2% | Claude Opus 5.5 |
| Toolathlon Verified Pass@3 | 82.4% | Coming soon | Coming soon |
| Toolathlon Verified Pass³ | 72.2% | Coming soon | Coming soon |
| Toolathlon Verified avg. turns | 26.9 turns | Coming soon | Coming soon |
| Terminal-Bench 2.0 | Coming soon | 88.3% | Coming soon |
| BrowseComp | Coming soon | 91.2% | Coming soon |
| DeepSearchQA | Coming soon | 95.0% | Coming soon |
| MCP Atlas | Coming soon | 84.2% | Coming soon |
| JobBench | Coming soon | 52.9% | Coming soon |
| APEX-Agents | Coming soon | 37.6% | Coming soon |
| SpreadsheetBench 2 | Coming soon | 34.8% | Coming soon |
| DECK-Bench | Coming soon | 73.5% | Coming soon |
| Terminal-Bench 2.1 (Vals) | Coming soon | 80.9% | Coming soon |
| ApprenticeBench | Coming soon | 18% | Coming soon |

### Coding

- Winner: Claude Opus 5.5
- Claude Opus 5.5 public-lane score: 80.1 (Supported · #2/159)
- Kimi K3 public-lane score: 66.5 (Supported · #12/159)

| Benchmark | Claude Opus 5.5 | Kimi K3 | Winner |
|-----------|-----------|-----------|--------|
| FrontierCode 1.1 Main | 54.4% | Coming soon | Coming soon |
| cursorBench40 | 57.8% | Coming soon | Coming soon |
| SWE-bench Pro | 89.9% | Coming soon | Coming soon |
| SWE Multilingual | 93.9% | Coming soon | Coming soon |
| SWE Multimodal | 61.4% | Coming soon | Coming soon |
| DeepSWE | 74.2% | 67.5% | Claude Opus 5.5 |
| FrontierCode 1.1 Extended | 63.6% | Coming soon | Coming soon |
| FrontierSWE v2 | 62.3% | 25.9% | Claude Opus 5.5 |
| ProgramBench | 91.2% | 77.8% | Claude Opus 5.5 |
| cursorBench32 | Coming soon | 60.8% | Coming soon |
| FrontierSWE | Coming soon | 81.2% | Coming soon |
| Kimi Code Bench v2 | Coming soon | 72.9% | Coming soon |
| sweMarathon | Coming soon | 42% | Coming soon |
| PostTrain Bench | Coming soon | 36.6% | Coming soon |
| MLS-Bench Lite | Coming soon | 48.3% | Coming soon |
| VulcanBench v3 | Coming soon | 73.7% | Coming soon |
| OpenHarmony Bench | Coming soon | 57.3% | Coming soon |
| LiveCodeBench (Vals) | Coming soon | 87.2% | Coming soon |
| SWE-bench (Vals) | Coming soon | 93.4% | Coming soon |

### Reasoning

- Winner: Coming soon
- Claude Opus 5.5 public-lane score: 78.5 (#4/17)
- Kimi K3 public-lane score: 78.5 (#3/17)

### Multimodal & Grounded

- Winner: Coming soon
- Claude Opus 5.5 public-lane score: 88.8 (#3/50)
- Kimi K3 public-lane score: 89.4 (#1/50)

| Benchmark | Claude Opus 5.5 | Kimi K3 | Winner |
|-----------|-----------|-----------|--------|
| Chartography (tools) | 89.0% | Coming soon | Coming soon |
| Chartography (no tools) | 64.4% | Coming soon | Coming soon |
| BenchCAD Vision2Code (no tools) | 0.730 | Coming soon | Coming soon |
| BenchCAD Vision2Code (tools) | 0.962 | Coming soon | Coming soon |
| Biomedical image analysis | 71.4% | Coming soon | Coming soon |
| OfficeQA | 78.9% | Coming soon | Coming soon |
| OfficeQA Pro | 67.7% | 63.3% | Claude Opus 5.5 |
| MMMU-Pro | Coming soon | 81.6% | Coming soon |
| MMMU-Pro w/ Python | Coming soon | 83.4% | Coming soon |
| CharXiv w/o tools | Coming soon | 84.8% | Coming soon |
| CharXiv | Coming soon | 91.3% | Coming soon |
| MathVision | Coming soon | 94.3% | Coming soon |
| MathVision w/ Python | Coming soon | 97.8% | Coming soon |
| BabyVision w/ Python | Coming soon | 85.7% | Coming soon |
| ZeroBench | Coming soon | 23.0% | Coming soon |
| ZeroBench w/ Python | Coming soon | 41.0% | Coming soon |
| WorldVQA ForceAnswer | Coming soon | 51.0% | Coming soon |
| OmniDocBench | Coming soon | 91.1% | Coming soon |
| PerceptionBench | Coming soon | 58.5% | Coming soon |

### Knowledge

- Winner: Coming soon
- Claude Opus 5.5 public-lane score: 81.8 (Estimated · #4/189)
- Kimi K3 public-lane score: 71.6 (Supported · #10/189)

| Benchmark | Claude Opus 5.5 | Kimi K3 | Winner |
|-----------|-----------|-----------|--------|
| HLE w/o tools | 64.4% | 43.5% | Claude Opus 5.5 |
| HealthBench (raw) | 68.1% | Coming soon | Coming soon |
| HealthBench (length-adjusted) | 60.6% | Coming soon | Coming soon |
| HealthBench Professional | 65.6% | Coming soon | Coming soon |
| HealthBench Professional (raw) | 77.1% | Coming soon | Coming soon |
| BioMysteryBench (human-solvable) | 89.3% | Coming soon | Coming soon |
| BioMysteryBench (human-difficult) | 50.0% | Coming soon | Coming soon |
| SpatialBench Verified | 72.0% | Coming soon | Coming soon |
| SingleCellBench | 61.2% | Coming soon | Coming soon |
| Morphology-to-molecule matching | 34.0% | Coming soon | Coming soon |
| Medicinal chemistry | 63.5% | Coming soon | Coming soon |
| Protein Design | 60.2% | Coming soon | Coming soon |
| Protein Design library ranking | 56.0% | Coming soon | Coming soon |
| De novo protein-binder design | 82.6% | Coming soon | Coming soon |
| Protocols (troubleshooting) | 73.7% | Coming soon | Coming soon |
| Protocols (understanding) | 69.0% | Coming soon | Coming soon |
| GPQA | Coming soon | 93.5% | Coming soon |
| GPQA-D | Coming soon | 93.5% | Coming soon |
| HLE | Coming soon | 56% | Coming soon |
| GPQA Diamond (Vals) | Coming soon | 92.9% | Coming soon |
| MMLU-Pro (Vals) | Coming soon | 88.0% | Coming soon |

### Multilingual

- Winner: Coming soon
- Claude Opus 5.5 public-lane score: Coming soon
- Kimi K3 public-lane score: Coming soon

| Benchmark | Claude Opus 5.5 | Kimi K3 | Winner |
|-----------|-----------|-----------|--------|
| GMMLU | 94.3% | Coming soon | Coming soon |
| MILU | 93.1% | Coming soon | Coming soon |

### Mathematics

- Winner: Coming soon
- Claude Opus 5.5 public-lane score: Coming soon
- Kimi K3 public-lane score: Coming soon

| Benchmark | Claude Opus 5.5 | Kimi K3 | Winner |
|-----------|-----------|-----------|--------|
| ArXivMath Aug. 2026 (no tools) | 91.2% | Coming soon | Coming soon |
| ArXivMath Aug. 2026 (tools) | 96.9% | Coming soon | Coming soon |

## FAQ

### Which is better overall, Claude Opus 5.5 or Kimi K3?

Claude Opus 5.5 is ahead overall on BenchLM right now.

### Where is the biggest gap between Claude Opus 5.5 and Kimi K3?

The widest category gap is in coding, where the averages are 80.1 for Claude Opus 5.5 and 66.5 for Kimi K3.

### How many benchmarks does Claude Opus 5.5 cover on BenchLM?

Claude Opus 5.5 currently has 47 sourced benchmark scores on BenchLM.

### How many benchmarks does Kimi K3 cover on BenchLM?

Kimi K3 currently has 44 sourced benchmark scores on BenchLM.

## Related Comparisons

- [Claude Opus 5.5 vs Claude Fable 5.1](/compare/claude-fable-5-1-vs-claude-opus-5-5)
- [Kimi K3 vs Claude Fable 5.1](/compare/claude-fable-5-1-vs-kimi-k3)
- [Claude Opus 5.5 vs GPT-6 Astra](/compare/claude-opus-5-5-vs-gpt-6-astra)
- [Kimi K3 vs GPT-6 Astra](/compare/gpt-6-astra-vs-kimi-k3)

## Explore More

- [Claude Opus 5.5 profile](/models/claude-opus-5-5)
- [Kimi K3 profile](/models/kimi-k3)
- [Compare Pricing](/llm-pricing)
- [Alternative Finder](/tools/alternative-finder)
- [LLM Selector](/tools/llm-selector)
- [Overall Rankings](/best/overall)
