Retrieval and Classification (Fastino report)
We show this table for reference; we do not rank on it.
Fastino reports GLiDE at 60.9 skill points on Retrieval and Classification in its Decision Index 0.2.1 run. The comparison quotes the published Jev 1.13.0 reference.
Published results and quoted references
The five area scores and overall index are chance-adjusted skill points. CLadder and CRUXEval are raw accuracies. Fastino reports a complete run, but publishes exact GLiDE values for only these eight metrics. Radar-chart gaps do not supply the other individual benchmark accuracies. We have not rerun the evaluation, and these results stay outside general model rankings. The launch chart also quotes five other overall references, including Liquid AI’s self-reported d1 reproduction; their category and task scores are not inferred.
| Metric | Unit | GLiDE | TypeSafe Jev 1.13.0 |
|---|---|---|---|
| Retrieval and Classification | Skill points | 60.9 | 55.4 |
A provider run with the official scorer
Fastino reports 64.81 skill points for GLiDE on Decision Index 0.2.1, compared with 57.91 for the published Jev 1.13.0 reference. Its September 30, 2026 announcement uses the official scorer; GLiDE is not listed on the public board.
The five area scores and overall index are chance-adjusted skill points. CLadder and CRUXEval are raw accuracies. Fastino reports a complete run, but publishes exact GLiDE values for only these eight metrics. Radar-chart gaps do not supply the other individual benchmark accuracies. We have not rerun the evaluation, and these results stay outside general model rankings. The launch chart also quotes five other overall references, including Liquid AI’s self-reported d1 reproduction; their category and task scores are not inferred.
Snapshot
About Retrieval and Classification (Fastino report)
Year
2026
Tasks
6 component benchmarks
Format
Chance-adjusted skill points
Difficulty
Varies by task and supplied answer options
The five area scores and overall index are chance-adjusted skill points. CLadder and CRUXEval are raw accuracies. Fastino reports a complete run, but publishes exact GLiDE values for only these eight metrics. Radar-chart gaps do not supply the other individual benchmark accuracies. We have not rerun the evaluation, and these results stay outside general model rankings. The area chart states RTX PRO 6000 hardware and 155,390 requests for the full suite; that count is not a sample size for every metric. The benchmark-owner artifact confirms the Jev reference identity.
Freshness and provenance
Version
Decision Index 0.2.1, Fastino run
Refresh cadence
Pinned provider report
Staleness state
Current
Question availability
Public benchmark kit; exact GLiDE outputs and complete task-level results not published in the launch
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
Is GLiDE on the public Decision Index leaderboard?
Fastino says GLiDE is not listed on the public Decision Index 0.2.1 leaderboard because submissions are paused. These values come from Fastino’s own run with the official scorer, published September 30, 2026. They are a provider report, not an independently reproduced leaderboard entry or a general model ranking.
Are skill points and benchmark accuracy interchangeable?
Decision Index skill points adjust results for the chance baseline before combining benchmarks. Raw accuracy is the percentage of correct answers under a particular task setup. The overall index and five areas use skill points; the published CLadder and CRUXEval figures use accuracy. Compare each metric only with the same protocol.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.