PostTrainBench v1.1
We show this table for reference; we do not rank on it.
Post-training four base language models across seven weighted benchmarks, with ten hours and one H100 per run.
Benchmark score on PostTrainBench v1.1 — October 7, 2026
We compile the PostTrainBench v1.1 rows from benchmark-owner or independent runs and provider self-reports. Claude Opus 5.5 leads the table at 49.3%, followed by Gemini 4 Argon (45.3%) and GPT-6 Astra (44.3%). We do not use these results to rank models overall.
Claude Opus 5.5
Anthropic
Gemini 4 Argon
GPT-6 Astra
OpenAI
14 modelsCodingCurrentDisplay onlyUpdated October 7, 2026
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Claude Opus 5.5Anthropic | 49.3% | Not reported | Closed |
| 2 | Gemini 4 ArgonGoogle | 45.3% | Not reported | Closed |
| 3 | GPT-6 AstraOpenAI | 44.3% | Not reported | Closed |
| 4 | Claude Fable 5.1Anthropic | 40.2% | Not reported | Closed |
| 5 | GPT-5.6 SolOpenAI | 36.2% | Not reported | Closed |
| 6 | Claude Opus 5Anthropic | 35.0% | Not reported | Closed |
| 7 | Claude Opus 4.8Anthropic | 32.9% | Not reported | Closed |
| 8 | Kimi K3Moonshot AI | 32.0% | Not reported | Pending |
| 9 | GLM-5.2Z.AI | 31.7% | Not reported | Open |
| 10 | Claude Opus 4.7Anthropic | 28.6% | Not reported | Closed |
| 11 | GPT-5.5OpenAI | 27.2% | Not reported | Closed |
| 12 | Grok 4.5xAI | 23.4% | Not reported | Closed |
| 13 | Gemini 3.1 ProGoogle | 22.0% | Not reported | Closed |
| 14 | GPT-5.4OpenAI | 19.0% | Not reported | Closed |
Among the reported PostTrainBench v1.1 rows, Claude Opus 5.5 is first at 49.3%. The third row is 5.0 points behind. The broader top-10 range is 20.7 points, so the table still separates the published systems.
14 models have been evaluated on PostTrainBench v1.1. The benchmark falls in the Coding category. PostTrainBench v1.1 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About PostTrainBench v1.1
Year
2026
Tasks
Post-training Qwen3 1.7B and 4B, SmolLM3 3B, and Gemma3 4B
Format
Weighted aggregate across four base models and seven benchmarks
Difficulty
Frontier agent evaluation
Version 1.1 audits contamination, external API use, prior-run lookup, and model identity. Flagged runs receive the base-model score. Native CLI leaderboard runs and Google OpenCode runs use different harnesses; each row records its source and setup.
Freshness and provenance
Version
PostTrainBench v1.1 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does PostTrainBench v1.1 measure?
Post-training four base language models across seven weighted benchmarks, with ten hours and one H100 per run.
Which model scores highest on PostTrainBench v1.1?
Claude Opus 5.5 by Anthropic currently leads with a score of 49.3% on PostTrainBench v1.1.
How many models are evaluated on PostTrainBench v1.1?
14 AI models have published results on PostTrainBench v1.1 in the BenchLM catalog.
Compare top models on PostTrainBench v1.1
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.