Skip to main content
BenchLM
Data

PostTrainBench v1.1

We show this table for reference; we do not rank on it.

Data verified 38 confirmed releases in the last 30 daysFollow model changes

Post-training four base language models across seven weighted benchmarks, with ten hours and one H100 per run.

Benchmark score on PostTrainBench v1.1 — October 7, 2026

We compile the PostTrainBench v1.1 rows from benchmark-owner or independent runs and provider self-reports. Claude Opus 5.5 leads the table at 49.3%, followed by Gemini 4 Argon (45.3%) and GPT-6 Astra (44.3%). We do not use these results to rank models overall.

14 modelsCodingCurrentDisplay onlyUpdated October 7, 2026

Benchmark score results for PostTrainBench v1.1
RankModel / configurationScoreParameters (B)Open / closed
1Claude Opus 5.5Anthropic
49.3%
Not reportedClosed
2Gemini 4 ArgonGoogle
45.3%
Not reportedClosed
3GPT-6 AstraOpenAI
44.3%
Not reportedClosed
4Claude Fable 5.1Anthropic
40.2%
Not reportedClosed
5GPT-5.6 SolOpenAI
36.2%
Not reportedClosed
6Claude Opus 5Anthropic
35.0%
Not reportedClosed
7Claude Opus 4.8Anthropic
32.9%
Not reportedClosed
8Kimi K3Moonshot AI
32.0%
Not reportedPending
9GLM-5.2Z.AI
31.7%
Not reportedOpen
10Claude Opus 4.7Anthropic
28.6%
Not reportedClosed
11GPT-5.5OpenAI
27.2%
Not reportedClosed
12Grok 4.5xAI
23.4%
Not reportedClosed
13Gemini 3.1 ProGoogle
22.0%
Not reportedClosed
14GPT-5.4OpenAI
19.0%
Not reportedClosed

Among the reported PostTrainBench v1.1 rows, Claude Opus 5.5 is first at 49.3%. The third row is 5.0 points behind. The broader top-10 range is 20.7 points, so the table still separates the published systems.

14 models have been evaluated on PostTrainBench v1.1. The benchmark falls in the Coding category. PostTrainBench v1.1 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About PostTrainBench v1.1

Year

2026

Tasks

Post-training Qwen3 1.7B and 4B, SmolLM3 3B, and Gemma3 4B

Format

Weighted aggregate across four base models and seven benchmarks

Difficulty

Frontier agent evaluation

Version 1.1 audits contamination, external API use, prior-run lookup, and model identity. Flagged runs receive the base-model score. Native CLI leaderboard runs and Google OpenCode runs use different harnesses; each row records its source and setup.

Freshness and provenance

Version

PostTrainBench v1.1 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does PostTrainBench v1.1 measure?

Post-training four base language models across seven weighted benchmarks, with ten hours and one H100 per run.

Which model scores highest on PostTrainBench v1.1?

Claude Opus 5.5 by Anthropic currently leads with a score of 49.3% on PostTrainBench v1.1.

How many models are evaluated on PostTrainBench v1.1?

14 AI models have published results on PostTrainBench v1.1 in the BenchLM catalog.

Last updated: October 7, 2026 · BenchLM version PostTrainBench v1.1 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.