Skip to main content
BenchLM

HealthBench length-adjusted score (HealthBench (length-adjusted))

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

HealthBench score after applying a verbosity penalty to model responses.

Benchmark score on HealthBench (length-adjusted) — September 27, 2026

We compile the HealthBench (length-adjusted) rows from provider self-reports. Claude Opus 5.5 leads the table at 60.6%, followed by GPT-6 Astra (58.3%) and Claude Opus 5 (57.8%). We do not use these results to rank models overall.

5 modelsKnowledgeCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (5 models)

Score
1
Claude Opus 5.5Anthropic · Closed
60.6%
2
GPT-6 AstraOpenAI · Closed
58.3%
3
Claude Opus 5Anthropic · Closed
57.8%
4
GPT-6 LunaOpenAI · Closed
54.5%
5
GPT-6 SolOpenAI · Closed
53.2%

Among the reported HealthBench (length-adjusted) rows, Claude Opus 5.5 is first at 60.6%. The third row is 2.8 points behind. The broader top-10 range is 7.4 points, so many of the published results sit in a relatively narrow band.

5 models have been evaluated on HealthBench (length-adjusted). The benchmark falls in the Knowledge category. HealthBench (length-adjusted) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About HealthBench (length-adjusted)

Year

2026

Tasks

5,000 multi-turn patient conversations

Format

Length-adjusted rubric score

Difficulty

Realistic healthcare conversations

Section 8.15.1 reports the length-adjusted result separately from the raw score. All Claude models used max effort, no tools, and five trials.

Freshness and provenance

Version

HealthBench (length-adjusted) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does HealthBench (length-adjusted) measure?

HealthBench score after applying a verbosity penalty to model responses.

Which model scores highest on HealthBench (length-adjusted)?

Claude Opus 5.5 by Anthropic currently leads with a score of 60.6% on HealthBench (length-adjusted).

How many models are evaluated on HealthBench (length-adjusted)?

5 AI models have been evaluated on HealthBench (length-adjusted) on BenchLM.

Last updated: September 27, 2026 · BenchLM version HealthBench (length-adjusted) 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.