# PostTrainBench v1.1

> Post-training four base language models across seven weighted benchmarks, with ten hours and one H100 per run.

Canonical page: https://benchlm.ai/benchmarks/posttrainbench-v1-1

- Category: [Coding](/coding)
- Last updated: October 7, 2026

## About PostTrainBench v1.1

- Year: 2026
- Tasks: Post-training Qwen3 1.7B and 4B, SmolLM3 3B, and Gemma3 4B
- Format: Weighted aggregate across four base models and seven benchmarks
- Difficulty: Frontier agent evaluation
- Paper: [PostTrainBench v1.1 leaderboard and methodology](https://posttrainbench.com/?version=v1.1)

Version 1.1 audits contamination, external API use, prior-run lookup, and model identity. Flagged runs receive the base-model score. Native CLI leaderboard runs and Google OpenCode runs use different harnesses; each row records its source and setup.

PostTrainBench v1.1 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (14 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 49.3% |
| 2 | [Gemini 4 Argon](/models/gemini-4-argon) | Google | 45.3% |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 44.3% |
| 4 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 40.2% |
| 5 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 36.2% |
| 6 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 35.0% |
| 7 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 32.9% |
| 8 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 32.0% |
| 9 | [GLM-5.2](/models/glm-5-2) | Z.AI | 31.7% |
| 10 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 28.6% |
| 11 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 27.2% |
| 12 | [Grok 4.5](/models/grok-4-5) | xAI | 23.4% |
| 13 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 22.0% |
| 14 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 19.0% |

## FAQ

### What does PostTrainBench v1.1 measure?

Post-training four base language models across seven weighted benchmarks, with ten hours and one H100 per run.

### Which model scores highest on PostTrainBench v1.1?

Claude Opus 5.5 by Anthropic currently leads with a score of 49.3% on PostTrainBench v1.1.

### How many models are evaluated on PostTrainBench v1.1?

14 AI models have published results on PostTrainBench v1.1 in the BenchLM catalog.

### Does PostTrainBench v1.1 affect BenchLM's overall score?

Not directly. PostTrainBench v1.1 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on PostTrainBench v1.1

- [Claude Opus 5.5 vs Gemini 4 Argon](/compare/claude-opus-5-5-vs-gemini-4-argon)
- [Gemini 4 Argon vs GPT-6 Astra](/compare/gemini-4-argon-vs-gpt-6-astra)
- [GPT-6 Astra vs Claude Fable 5.1](/compare/claude-fable-5-1-vs-gpt-6-astra)
- [Claude Fable 5.1 vs GPT-5.6 Sol](/compare/claude-fable-5-1-vs-gpt-5-6-sol)
