# HealthBench Hard

> A harder subset of OpenAI's HealthBench for evaluating open-ended medical and health reasoning with rubric-based grading.

Canonical page: https://benchlm.ai/benchmarks/healthbench-hard

- Category: [Knowledge](/knowledge)
- Last updated: September 27, 2026

## About HealthBench Hard

- Year: 2026
- Tasks: 1,000 health prompts
- Format: Open-ended health evaluation
- Difficulty: Advanced health reasoning
- Paper: [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology)

Meta describes HealthBench Hard as a 1,000-prompt subset of OpenAI's HealthBench, graded with the same simple-evals implementation and a GPT-4.1-based judge. BenchLM treats it as a display-only health benchmark reference.

HealthBench Hard is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (11 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Muse Spark](/models/muse-spark) | Meta | 42.8% |
| 2 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 40.1% |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 36.6% |
| 4 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 33.1% |
| 5 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 32.7% |
| 6 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 32.0% |
| 7 | [GPT-6 Luna](/models/gpt-6-luna) | OpenAI | 31.4% |
| 8 | [GPT-6 Sol](/models/gpt-6-sol) | OpenAI | 30.1% |
| 9 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 20.6% |
| 10 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 20.3% |
| 11 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 14.8% |

## FAQ

### What does HealthBench Hard measure?

A harder subset of OpenAI's HealthBench for evaluating open-ended medical and health reasoning with rubric-based grading.

### Which model scores highest on HealthBench Hard?

Muse Spark by Meta currently leads with a score of 42.8% on HealthBench Hard.

### How many models are evaluated on HealthBench Hard?

11 AI models have been evaluated on HealthBench Hard on BenchLM.

### Does HealthBench Hard affect BenchLM's overall score?

Not directly. HealthBench Hard is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on HealthBench Hard

- [Muse Spark vs GPT-5.4](/compare/gpt-5-4-vs-muse-spark)
- [GPT-5.4 vs GPT-6 Astra](/compare/gpt-5-4-vs-gpt-6-astra)
- [GPT-6 Astra vs GPT-5.6 Sol](/compare/gpt-5-6-sol-vs-gpt-6-astra)
- [GPT-5.6 Sol vs GPT-5.6 Terra](/compare/gpt-5-6-sol-vs-gpt-5-6-terra)
