# PostTrain Bench

> A software-engineering benchmark for post-training infrastructure and implementation tasks, evaluated through the official Harbor implementation.

Canonical page: https://benchlm.ai/benchmarks/posttrainbench

- Category: [Coding](/coding)
- Last updated: September 27, 2026

## About PostTrain Bench

- Year: 2026
- Tasks: Post-training software-engineering tasks
- Format: Harbor agent evaluation
- Difficulty: Frontier software engineering
- Paper: [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3)

Moonshot evaluates Kimi K3 at maximum reasoning effort with the official Harbor implementation and Claude Code harness, averaged over three runs. BenchLM stores the provider-exact value as display-only launch evidence.

PostTrain Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GLM-5.3](/models/glm-5-3) | Z.AI | 39.8% |
| 2 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 36.6% |
| 3 | [Hy4 preview](/models/hy4-preview) | Tencent | 35.6% |

## FAQ

### What does PostTrain Bench measure?

A software-engineering benchmark for post-training infrastructure and implementation tasks, evaluated through the official Harbor implementation.

### Which model scores highest on PostTrain Bench?

GLM-5.3 by Z.AI currently leads with a score of 39.8% on PostTrain Bench.

### How many models are evaluated on PostTrain Bench?

3 AI models have been evaluated on PostTrain Bench on BenchLM.

### Does PostTrain Bench affect BenchLM's overall score?

Not directly. PostTrain Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on PostTrain Bench

- [GLM-5.3 vs Kimi K3](/compare/glm-5-3-vs-kimi-k3)
- [Kimi K3 vs Hy4 preview](/compare/hy4-preview-vs-kimi-k3)
