# Scale Labs HiL-Bench (HiL-Bench)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-hil

- Category: [Agentic](/agentic)
- Last updated: September 29, 2026 snapshot

## About HiL-Bench

- Year: 2026
- Tasks: 17 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/hil)

BenchLM mirrors 17 published rows from the HiL-Bench public table captured on September 29, 2026 snapshot.

HiL-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (17 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 61.50% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 57.00% |
| 3 | [Claude Fable 5](/models/claude-fable) | Anthropic | 56.33% |
| 4 | [GLM-5.2](/models/glm-5-2) | Z.AI | 43.67% |
| 5 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 41.67% |
| 6 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 41.47% |
| 7 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 39.67% |
| 8 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 38.33% |
| 9 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 35.33% |
| 10 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 35.33% |
| 11 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 32.33% |
| 12 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 27.67% |
| 13 | [Grok-4.20](https://labs.scale.com/leaderboard/hil) | xAI | 20.00% |
| 14 | [Kimi-k2.6](https://labs.scale.com/leaderboard/hil) | Moonshot AI | 18.67% |
| 15 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 9.67% |
| 16 | [MiniMax M2.5](/models/minimax-m2-5) | MiniMax | 6.33% |
| 17 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 4.33% |

## FAQ

### What does HiL-Bench measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published HiL-Bench snapshot?

Claude Fable 5.1 currently leads the published HiL-Bench snapshot with a score of 61.50%.

### How many models are evaluated on HiL-Bench?

The September 29, 2026 snapshot contains 17 AI models.

### Does HiL-Bench affect BenchLM's overall score?

Not directly. HiL-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
