Skip to main content
BenchLM

Scale Labs HiL-Bench (HiL-Bench)

We show this table for reference; we do not rank on it.

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Scale score on HiL-Bench — September 29, 2026 snapshot

We mirror the published scale score view for HiL-Bench. Claude Fable 5.1 leads the public snapshot at 61.5%, followed by Claude Opus 5 (57%) and Claude Fable 5 (56.3%). We do not use these results to rank models overall.

17 modelsAgenticCurrentDisplay onlyUpdated September 29, 2026 snapshot

Scale score table (17 models)

Score
1
Claude Fable 5.1Anthropic · Closed
61.5%
2
Claude Opus 5Anthropic · Closed
57%
3
Claude Fable 5Anthropic · Closed
56.3%
4
GLM-5.2Z.AI · Open weight
43.7%
5
Claude Opus 4.7Anthropic · Closed
41.7%
6
Gemini 3.8 FlashGoogle · Closed
41.5%
7
GPT-5.5OpenAI · Closed
39.7%
8
Claude Opus 4.6Anthropic · Closed
38.3%
9
Claude Opus 4.8Anthropic · Closed
35.3%
10
Gemini 3.1 ProGoogle · Closed
35.3%
11
GPT-5.6 SolOpenAI · Closed
32.3%
12
Gemini 3.5 FlashGoogle · Closed
27.7%
13
20%
14
Kimi-k2.6Moonshot AI
18.7%
15
GPT-5.4OpenAI · Closed
9.7%
16
MiniMax M2.5MiniMax · Closed
6.3%
17
GPT-5.3 CodexOpenAI · Closed
4.3%

How to read this leaderboard

Compare the published configurations as complete evaluation systems. The source can combine a base model, agent scaffold, tools, budget, and inference setting in each result.

Operator receipt: 17 sourced rows are currently displayable on this page; the leading published row is Claude Fable 5.1 at 61.5%.

Honest limit: This Scale table is display-only context, not benchmark provenance or a weighted model-only comparison.

How BenchLM shows HiL-Bench

BenchLM mirrors 17 published rows from Scale Labs’ public HiL-Bench leaderboard, captured on September 29, 2026 snapshot.

The table is display only. It is useful context for a published agent or model configuration, but it does not enter BenchLM’s overall or category rankings.

Snapshot

17 published rowsScale Labs sourceDisplay only

The published HiL-Bench snapshot places Claude Fable 5.1 first at 61.5%. The third row is 5.2 points behind. The broader top-10 range is 26.2 points, so the table still separates the published systems.

17 models have been evaluated on HiL-Bench. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. HiL-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About HiL-Bench

Year

2026

Tasks

17 published rows

Format

Published Scale leaderboard score

Difficulty

External agent and model evaluation

BenchLM mirrors 17 published rows from the HiL-Bench public table captured on September 29, 2026 snapshot.

Freshness and provenance

Version

HiL-Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does HiL-Bench measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Which model leads the published HiL-Bench snapshot?

Claude Fable 5.1 currently leads the published HiL-Bench snapshot with 61.5% scale score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on HiL-Bench?

The September 29, 2026 snapshot snapshot contains 17 AI models.

Last updated: September 29, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.