# Kimi K3 vs BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI)

> Side-by-side benchmark comparison for Kimi K3 and BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) across agentic, coding, multimodal, knowledge, reasoning, multilingual, and math tasks.

- Canonical page: https://benchlm.ai/compare/kimi-k3-vs-native-contractnli-bert-base-fine-tuned-on-case-law-and-contract-corpora-chalkidis-et-al-2020
- Last updated: October 1, 2026

- Shared sourced benchmarks: 0
- HTML indexing: indexable
- Ranking lane: BenchAlign v5.8

## Quick Verdict

No benchmark verdict yet: there is no shared sourced benchmark row for Kimi K3 and BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI).

## Summary

- Kimi K3 and BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) do not currently share a sourced benchmark row. The mirror therefore compares metadata, pricing, context windows, and each model's separately reported coverage without naming a benchmark winner.

## Model Snapshot

| Property | Kimi K3 | BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) |
|----------|----------|----------|
| Creator | Moonshot AI | Koreeda and Manning |
| Type | Pending | Research system |
| Reasoning | Reasoning | Unspecified |
| Context | 1.05M | null |
| Independent public score (not head-to-head) | 72.12 | null |
| Benchmarks Covered | 48 | 0 |
| Pricing (input/output) | $3.00 / $15.00 | Pricing unavailable / Pricing unavailable |

## Category Breakdown

### Agentic

- Winner: Insufficient shared evidence
- Kimi K3 public-lane score: 67.8 (Supported · #10/119)
- BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) public-lane score: Coming soon

| Benchmark | Kimi K3 | BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) | Winner |
|-----------|-----------|-----------|--------|
| Terminal-Bench 2.1 | 88.3% | Coming soon | Coming soon |
| BrowseComp | 91.2% | Coming soon | Coming soon |
| DeepSearchQA | 95.0% | Coming soon | Coming soon |
| Toolathlon-Verified | 73.2% | Coming soon | Coming soon |
| MCP Atlas | 84.2% | Coming soon | Coming soon |
| AutomationBench | 30.8% | Coming soon | Coming soon |
| JobBench | 52.9% | Coming soon | Coming soon |
| APEX-Agents | 37.6% | Coming soon | Coming soon |
| SpreadsheetBench 2 | 34.8% | Coming soon | Coming soon |
| DECK-Bench | 73.5% | Coming soon | Coming soon |
| Terminal-Bench 2.1 (Vals) | 80.9% | Coming soon | Coming soon |
| ApprenticeBench | 18% | Coming soon | Coming soon |

### Coding

- Winner: Insufficient shared evidence
- Kimi K3 public-lane score: 61.4 (Supported · #18/144)
- BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) public-lane score: Coming soon

| Benchmark | Kimi K3 | BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) | Winner |
|-----------|-----------|-----------|--------|
| DeepSWE | 67.5% | Coming soon | Coming soon |
| CursorBench 3.2 | 60.8% | Coming soon | Coming soon |
| FrontierSWE | 81.2% | Coming soon | Coming soon |
| ProgramBench | 77.8% | Coming soon | Coming soon |
| Kimi Code Bench v2 | 72.9% | Coming soon | Coming soon |
| sweMarathon | 42% | Coming soon | Coming soon |
| PostTrain Bench | 36.6% | Coming soon | Coming soon |
| MLS-Bench Lite | 48.3% | Coming soon | Coming soon |
| VulcanBench v3 | 73.7% | Coming soon | Coming soon |
| OpenHarmony Bench | 57.3% | Coming soon | Coming soon |
| FrontierSWE v2 | 25.9% | Coming soon | Coming soon |
| LiveCodeBench (Vals) | 87.2% | Coming soon | Coming soon |
| SWE-bench (Vals) | 93.4% | Coming soon | Coming soon |
| PostTrainBench v1.1 | 32.0% | Coming soon | Coming soon |

### Reasoning

- Winner: Insufficient shared evidence
- Kimi K3 public-lane score: 65.8 (#18/27)
- BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) public-lane score: Coming soon

| Benchmark | Kimi K3 | BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) | Winner |
|-----------|-----------|-----------|--------|
| ARC-AGI-1 | 94.50% | Coming soon | Coming soon |
| ARC-AGI-2 | 60.4% | Coming soon | Coming soon |

### Multimodal & Grounded

- Winner: Insufficient shared evidence
- Kimi K3 public-lane score: 89.4 (#1/49)
- BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) public-lane score: Coming soon

| Benchmark | Kimi K3 | BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) | Winner |
|-----------|-----------|-----------|--------|
| OfficeQA Pro | 63.3% | Coming soon | Coming soon |
| MMMU-Pro | 81.6% | Coming soon | Coming soon |
| MMMU-Pro w/ Python | 83.4% | Coming soon | Coming soon |
| CharXiv w/o tools | 84.8% | Coming soon | Coming soon |
| CharXiv | 91.3% | Coming soon | Coming soon |
| MathVision | 94.3% | Coming soon | Coming soon |
| MathVision w/ Python | 97.8% | Coming soon | Coming soon |
| BabyVision w/ Python | 85.7% | Coming soon | Coming soon |
| ZeroBench | 23.0% | Coming soon | Coming soon |
| ZeroBench w/ Python | 41.0% | Coming soon | Coming soon |
| WorldVQA ForceAnswer | 51.0% | Coming soon | Coming soon |
| OmniDocBench | 91.1% | Coming soon | Coming soon |
| PerceptionBench | 58.5% | Coming soon | Coming soon |

### Knowledge

- Winner: Insufficient shared evidence
- Kimi K3 public-lane score: 67.8 (Supported · #18/170)
- BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) public-lane score: Coming soon

| Benchmark | Kimi K3 | BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) | Winner |
|-----------|-----------|-----------|--------|
| GPQA | 93.5% | Coming soon | Coming soon |
| GPQA-D | 93.5% | Coming soon | Coming soon |
| HLE | 56% | Coming soon | Coming soon |
| HLE w/o tools | 43.5% | Coming soon | Coming soon |
| GPQA Diamond (Vals) | 92.9% | Coming soon | Coming soon |
| MMLU-Pro (Vals) | 88.0% | Coming soon | Coming soon |

### Instruction Following

- Winner: Insufficient shared evidence
- Kimi K3 public-lane score: Coming soon
- BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) public-lane score: Coming soon

| Benchmark | Kimi K3 | BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) | Winner |
|-----------|-----------|-----------|--------|
| Gray Swan IPI (15 attempts) | 52.7% | Coming soon | Coming soon |

## FAQ

### Can I compare Kimi K3 and BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) on BenchLM yet?

Not fully yet. BenchLM is tracking both models, but sourced benchmark coverage is still incomplete for a fair score-level comparison.

### Why does this page show "coming soon" values?

BenchLM only calls winners when public benchmark coverage is available on both sides of the comparison.

## Related Comparisons

- [Kimi K3 vs GPT-6 Astra](/compare/gpt-6-astra-vs-kimi-k3)
- [BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) vs GPT-6 Astra](/compare/gpt-6-astra-vs-native-contractnli-bert-base-fine-tuned-on-case-law-and-contract-corpora-chalkidis-et-al-2020)
- [Kimi K3 vs Claude Opus 5.5](/compare/claude-opus-5-5-vs-kimi-k3)
- [BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) vs Claude Opus 5.5](/compare/claude-opus-5-5-vs-native-contractnli-bert-base-fine-tuned-on-case-law-and-contract-corpora-chalkidis-et-al-2020)

## Explore More

- [Kimi K3 profile](/models/kimi-k3)
- [BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI) profile](/models/native-contractnli-bert-base-fine-tuned-on-case-law-and-contract-corpora-chalkidis-et-al-2020)
- [Compare Pricing](/llm-pricing)
- [Alternative Finder](/tools/alternative-finder)
- [LLM Selector](/tools/llm-selector)
- [Overall Rankings](/best/overall)
