# Kimi K3 Benchmark Scores & Performance

> Kimi K3 by Moonshot AI scores 73.99/100 overall, ranking #10 out of 505 AI models.

Canonical page: https://benchlm.ai/models/kimi-k3

Last updated: September 22, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Moonshot AI |
| Source Type | Pending |
| Reasoning Type | Reasoning |
| Context Window | 1.05M |
| Overall Score | 73.99/100 |
| Overall Rank | #10 of 505 |

## Family & Coverage

- Family: Kimi K3
- Variant: base
- Benchmarks covered: 73 of 481
- Related earlier model: [Kimi K2.6](/models/kimi-2-6)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 88.3% |
| [BrowseComp](/benchmarks/browsecomp) | 91.2% |
| [DeepSearchQA](/benchmarks/deepsearchqa) | 95.0% |
| [Toolathlon-Verified](/benchmarks/toolathlonverified) | 73.2% |
| [MCP Atlas](/benchmarks/mcpatlas) | 84.2% |
| [AutomationBench](/benchmarks/automationbench) | 30.8% |
| [JobBench](/benchmarks/jobbench) | 52.9% |
| [APEX-Agents](/benchmarks/apexagents) | 37.6% |
| [SpreadsheetBench 2](/benchmarks/spreadsheetbench2) | 34.8% |
| [DECK-Bench](/benchmarks/deckbench) | 73.5% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 51.2% |
| [AA Briefcase](/benchmarks/aabriefcaseelo) | 1510 |
| [AA AutomationBench](/benchmarks/aaautomationbench) | 58.3% |
| [AA EnterpriseOps-Gym](/benchmarks/aaenterpriseopsgym) | 45.3% |
| [AA Harvey LAB](/benchmarks/aaharveylab) | 94.6% |
| [AA Tau3 Banking](/benchmarks/aatau3banking) | 46.0% |
| [APEX-Agents-AA](/benchmarks/apexagentsaa) | 41.3% |
| [AA ITBench](/benchmarks/aaitbench) | 47.7% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 80.9% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1524 |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 50.6% |
| [GDP.pdf](/benchmarks/aagdppdf) | 22.0% |
| [AA-AnalystAgent](/benchmarks/aaanalystagent) | 38.8% |
| [ApprenticeBench](/benchmarks/apprenticebench) | 18% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [DeepSWE](/benchmarks/deepswe) | 67.5% |
| [CursorBench 3.2](/benchmarks/cursorbench32) | 60.8% |
| [FrontierSWE](/benchmarks/frontierswe) | 81.2% |
| [ProgramBench](/benchmarks/programbench) | 77.8% |
| [Kimi Code Bench v2](/benchmarks/kimicodebenchv2) | 72.9% |
| [sweMarathon](/benchmarks/swemarathon) | 42% |
| [PostTrain Bench](/benchmarks/posttrainbench) | 36.6% |
| [MLS-Bench Lite](/benchmarks/mlsbenchlite) | 48.3% |
| [AA-SciCode](/benchmarks/aascicode) | 59.5% |
| [VulcanBench v3](/benchmarks/vulcanbench) | 73.7% |
| [OpenHarmony Bench](/benchmarks/openharmonybench) | 57.3% |
| [FrontierSWE v2](/benchmarks/frontierswev2) | 25.9% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 87.2% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 93.4% |
| [AA Coding Index](/benchmarks/aacodingindex) | 76.2% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [OfficeQA Pro](/benchmarks/officeqapro) | 63.3% |
| [MMMU-Pro](/benchmarks/mmmu-pro) | 81.6% |
| [MMMU-Pro w/ Python](/benchmarks/mmmupropython) | 83.4% |
| [CharXiv w/o tools](/benchmarks/charxivnotools) | 84.8% |
| [CharXiv](/benchmarks/charxiv) | 91.3% |
| [MathVision](/benchmarks/mathvision) | 94.3% |
| [MathVision w/ Python](/benchmarks/mathvisionpython) | 97.8% |
| [BabyVision w/ Python](/benchmarks/babyvisionpython) | 85.7% |
| [ZeroBench](/benchmarks/zerobench) | 23.0% |
| [ZeroBench w/ Python](/benchmarks/zerobenchpython) | 41.0% |
| [WorldVQA ForceAnswer](/benchmarks/worldvqaforceanswer) | 51.0% |
| [OmniDocBench](/benchmarks/omnidocbench) | 91.1% |
| [PerceptionBench](/benchmarks/perceptionbench) | 58.5% |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 80.5% |
| [Design Arena Website](/benchmarks/designarenawebsite) | 1350 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 88.7% |
| [CritPt](/benchmarks/critpt) | 23.4% |
| [MLCR-AA](/benchmarks/aamlcr) | 38.3% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA](/benchmarks/gpqa) | 93.5% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 93.5% |
| [HLE](/benchmarks/hle) | 56% |
| [HLE w/o tools](/benchmarks/hlenotools) | 43.5% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 43.6% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 93.5% |
| [AA-HLE](/benchmarks/aahle) | 46.9% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 19.7% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 47.6% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 53.2% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 92.9% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 88.0% |

## Other Moonshot AI Models

- [Kimi K2.7 Code](/models/kimi-k2-7-code) - Score: 66.23
- [Kimi K2.6](/models/kimi-2-6) - Score: 66.13
- [Kimi K2.5 (Reasoning)](/models/kimi-k2-5-reasoning) - Score: 57.88
- [Kimi K2.5](/models/kimi-k2-5) - Score: 54.41
- [Moonshot v1](/models/moonshot-v1) - Score: 43.57
- [Kimi K2](/models/kimi-k2) - Score: 23.36
- [Kimi-Audio 7B](/models/kimi-audio-7b) - Score: not computed
