Skip to main content
BenchLM

Decision Index 0.2.1 (Fastino report)

We show this table for reference; we do not rank on it.

Data verified 43 confirmed releases in the last 30 daysFollow model changes

Fastino reports GLiDE at 64.81 skill points on Decision Index 0.2.1 in its Decision Index 0.2.1 run. The comparison quotes the published Jev 1.13.0 reference.

Published results and quoted references

The five area scores and overall index are chance-adjusted skill points. CLadder and CRUXEval are raw accuracies. Fastino reports a complete run, but publishes exact GLiDE values for only these eight metrics. Radar-chart gaps do not supply the other individual benchmark accuracies. We have not rerun the evaluation, and these results stay outside general model rankings. The launch chart also quotes five other overall references, including Liquid AI’s self-reported d1 reproduction; their category and task scores are not inferred.

Fastino-reported GLiDE results and published Jev reference on Decision Index 0.2.1
MetricUnitGLiDETypeSafe Jev 1.13.0
Decision Index 0.2.1Skill points64.8157.91

Other overall references in Fastino’s launch chart

Additional overall skill scores quoted in the Fastino GLiDE launch chart
Evaluated systemSkill pointsSource qualification
d158.9Liquid AI self-reported reproduction, quoted by Fastino
Surogate Rune 26B-A4B v357.44Published Decision Index reference quoted by Fastino
Decider chat · Gemma-4-31B57.33Inference technique on Gemma-4-31B; published reference quoted by Fastino
AutoJev-27B56.40Published Decision Index reference quoted by Fastino
simple-jev · Qwen3.8-27B55.74Featherless inference technique; published reference quoted by Fastino

A provider run with the official scorer

Fastino reports 64.81 skill points for GLiDE on Decision Index 0.2.1, compared with 57.91 for the published Jev 1.13.0 reference. Its September 30, 2026 announcement uses the official scorer; GLiDE is not listed on the public board.

The five area scores and overall index are chance-adjusted skill points. CLadder and CRUXEval are raw accuracies. Fastino reports a complete run, but publishes exact GLiDE values for only these eight metrics. Radar-chart gaps do not supply the other individual benchmark accuracies. We have not rerun the evaluation, and these results stay outside general model rankings. The launch chart also quotes five other overall references, including Liquid AI’s self-reported d1 reproduction; their category and task scores are not inferred.

About Decision Index 0.2.1 (Fastino report)

Year

2026

Tasks

38 component benchmarks

Format

Chance-adjusted skill points

Difficulty

Varies by task and supplied answer options

The five area scores and overall index are chance-adjusted skill points. CLadder and CRUXEval are raw accuracies. Fastino reports a complete run, but publishes exact GLiDE values for only these eight metrics. Radar-chart gaps do not supply the other individual benchmark accuracies. We have not rerun the evaluation, and these results stay outside general model rankings. The area chart states RTX PRO 6000 hardware and 155,390 requests for the full suite; that count is not a sample size for every metric. The benchmark-owner artifact confirms the Jev reference identity.

Freshness and provenance

Version

Decision Index 0.2.1, Fastino run

Refresh cadence

Pinned provider report

Staleness state

Current

Question availability

Public benchmark kit; exact GLiDE outputs and complete task-level results not published in the launch

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

Is GLiDE on the public Decision Index leaderboard?

Fastino says GLiDE is not listed on the public Decision Index 0.2.1 leaderboard because submissions are paused. These values come from Fastino’s own run with the official scorer, published September 30, 2026. They are a provider report, not an independently reproduced leaderboard entry or a general model ranking.

Are skill points and benchmark accuracy interchangeable?

Decision Index skill points adjust results for the chance baseline before combining benchmarks. Raw accuracy is the percentage of correct answers under a particular task setup. The overall index and five areas use skill points; the published CLadder and CRUXEval figures use accuracy. Compare each metric only with the same protocol.

Last updated: September 30, 2026 · Fastino report with separately qualified reference scores

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.