Skip to main content
BenchLM

Decision Models

JevBench v1.5.4 · retrieved September 30, 2026 · 112 systems · 106 ranked
Compare AI systems that choose an allowed answer and report probabilities for routing, classification, and scoring. JevBench keeps decision quality, confidence, latency, and cost visible so you can check the trade-offs for your application.

Typed decisions your code can use

Decision models take application state, a written question, and an allowed answer space. Their output fits a check your code can act on. A bounded answer can still be wrong, and reported confidence needs testing on the decisions your application actually makes.

Decision primitives and example tasks
PrimitiveExample taskOutput
ChoiceRoute a support ticket to billing, technical support, or account support.An option and its probability distribution.
NoulCheck whether a request meets a written approval rule.A probability that a yes/no proposition holds.
ScoreRate a bug against an ordered severity rubric.A score and probabilities over the allowed levels.

The OpenRouter Jev explainer shows each primitive with a real API response.

Perplexity Decider v1 27B

Perplexity released Decider v1 27B on October 1, 2026, as a Qwen3.8-27B fine-tune with Apache 2.0 weights. It reads text, JSON, and images, then returns yes/no probabilities, choices, or rubric scores. The hosted Decisions API charges $0.04 per million input tokens, with free output and no per-request fee.

GLiDE makes structured decisions with adaptive reasoning

Fastino released GLiDE on September 30, 2026. It returns Noul, Choice, and Score answers with probabilities and confidence, and spends additional reasoning on uncertain decisions. The API accepts 40,000 tokens per rendered question prompt. Fastino lists $0.30 per million input tokens and free output; input usage sums internal passes across questions.

GLiDE on Decision Index 0.2.1

Fastino reports 64.81 skill points for GLiDE on Decision Index 0.2.1, compared with 57.91 for the published Jev 1.13.0 reference. Its September 30, 2026 announcement uses the official scorer; GLiDE is not listed on the public board.

Fastino-reported GLiDE results and published Jev reference on Decision Index 0.2.1
MetricUnitGLiDETypeSafe Jev 1.13.0
Decision Index 0.2.1Skill points64.8157.91
Knowledge and ReasoningSkill points62.951.4
Language UnderstandingSkill points63.862.0
Retrieval and ClassificationSkill points60.955.4
Tools and AutomationSkill points83.575.1
Arts and Human TasteSkill points46.037.7
CLadderAccuracy88.7%72.6%
CRUXEvalAccuracy92.6%73.0%

Other overall references in Fastino’s launch chart

Additional overall skill scores quoted in the Fastino GLiDE launch chart
Evaluated systemSkill pointsSource qualification
d158.9Liquid AI self-reported reproduction, quoted by Fastino
Surogate Rune 26B-A4B v357.44Published Decision Index reference quoted by Fastino
Decider chat · Gemma-4-31B57.33Inference technique on Gemma-4-31B; published reference quoted by Fastino
AutoJev-27B56.40Published Decision Index reference quoted by Fastino
simple-jev · Qwen3.8-27B55.74Featherless inference technique; published reference quoted by Fastino

The five area scores and overall index are chance-adjusted skill points. CLadder and CRUXEval are raw accuracies. Fastino reports a complete run, but publishes exact GLiDE values for only these eight metrics. Radar-chart gaps do not supply the other individual benchmark accuracies. We have not rerun the evaluation, and these results stay outside general model rankings. The launch chart also quotes five other overall references, including Liquid AI’s self-reported d1 reproduction; their category and task scores are not inferred.

Decision Index report and limits · Fastino launch and charts

Perplexity’s 11-task decision panel

Perplexity reports 85.71% accuracy for Decider v1 27B, 84.51% for TypeSafe Jev 1.13.0, and 74.76% for Qwen3.8-27B on a fixed 7,210-row panel. The chart covers September 2026 and was published October 1. Decider was measured through the Perplexity API.

These are provider-reported results. The task samples have different sizes, so the published overall weights them differently. Prompts, sampled item identities, label mappings, baseline inference settings, and uncertainty are not specified in the model card. This panel stays separate from full-dataset results and general model rankings; JevBench public-hard accuracy is separate from its full composite.

Perplexity-reported accuracy on the September 2026 decision panel
BenchmarkRowsJev 1.13.0Qwen3.8-27BPerplexity Decider
WinoGrande1,00090.70%73.10%83.30%
FinancialPhraseBank99976.98%75.68%84.18%
RAGTruth1,50077.27%61.53%88.80%
JudgeBench35078.57%68.86%78.29%
BBH75094.27%72.80%82.80%
JevBench public hard10173.27%72.28%70.30%
TabFact50089.80%78.60%90.60%
ContractNLI51077.45%80.78%80.78%
Circa50084.60%87.00%89.20%
Belebele50095.00%93.20%94.00%
TruthfulQA binary50092.00%82.80%85.40%
Overall7,21084.51%74.76%85.71%

Emphasized values mark the best percentage in each row, including ties. Read the panel result notes, the pinned Perplexity model card, or the official launch and chart.

JevBench decision-model leaderboard

JevBench v1.5.4 covers 1624 decisions. Its official score combines Intelligence, Calibration, Speed, and Cost; it is an index, rather than percent accuracy. The 112 configurations include 106 ranked systems and retain the source's unranked entries.

The headline uses option A, selected after the method owner reviewed the results. Published composite intervals and adjacent-pair statistical ties remain attached to each configuration. Cost per 1,000 decisions retains source tariffs or estimates. This composite stays outside general model rankings.

Option A · joint leaders (statistical tie). The ≈ marker identifies a published statistical tie with the next row. Missing pairwise markers establish neither a tie nor a separation; 95% intervals below describe individual composite scores.

Official JevBench 1.5 results (112 systems)
RankSystemScoreIntelligenceCalibrationSpeedCostSealed IntelligenceUSD / 1,000 decisions
1 ≈Cygnetblockbrain · system-one-open
Evaluation setup

Source entry: Cygnet (blockbrain, frozen Gemma-4-12B-it). Published attribution: blockbrain.

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.230 s; p95 0.348 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 73.16; C 73.70.

choice: open 82.92; sealed 80.29 (chance-corrected competence).

noul: open 57.86; sealed 55.67 (chance-corrected competence).

score: open 76.60; sealed 73.20 (chance-corrected competence).

73.7095% CI 72.36–74.4671.0987.0190.9756.4369.720.02834estimate
2Winnow-12B Q8Eldan Ring · jev-rebuild
Evaluation setup

Source entry: Winnow-12B Q8. Published attribution: Eldan Ring.

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.340 s; p95 0.715 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 73.47; C 73.23.

choice: open 77.61; sealed 82.04 (chance-corrected competence).

noul: open 58.61; sealed 66.15 (chance-corrected competence).

score: open 81.21; sealed 80.99 (chance-corrected competence).

73.2395% CI 72.02–73.9974.4384.0786.1456.5676.390.02806estimate
3Jev 1.13.0TypeSafe AI · jevAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Jev 1.13.0 (TypeSafe AI). Published attribution: TypeSafe AI.

Model author or lab

the operator's hosted API

Adjusted latency: p50 0.616 s; p95 0.674 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1); no exact base-model floor applies

Secondary composites: B 72.11; C 72.13.

choice: open 85.73; sealed 87.59 (chance-corrected competence).

noul: open 47.75; sealed 48.62 (chance-corrected competence).

score: open 81.16; sealed 81.14 (chance-corrected competence).

72.1395% CI 71.01–72.6172.0088.0383.8154.7372.450.03230estimate
4JevK5 v0.3 4Ballebee · unclassified
Evaluation setup

Source entry: JevK5 v0.3 (4B). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.182 s; p95 0.243 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 68.11; C 63.24.

choice: open 80.25; sealed 74.25 (chance-corrected competence).

noul: open 42.33; sealed 14.98 (chance-corrected competence).

score: open 62.22; sealed 63.59 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

71.9095% CI 69.39–72.9556.2788.3493.5563.0750.940.01702estimate
5Plumb-4Bcrh225 · unclassified
Evaluation setup

Source entry: Plumb-4B (crh225, JevK5 v0.2 + LoRA). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.185 s; p95 0.245 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 67.75; C 62.00.

choice: open 80.43; sealed 73.30 (chance-corrected competence).

noul: open 42.95; sealed 16.58 (chance-corrected competence).

score: open 59.26; sealed 62.57 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

71.5695% CI 69.20–72.7355.8587.4493.4563.0750.810.01702estimate
6 ≈Jev-Omniakhilaaa3 · jev-rebuild
Evaluation setup

Source entry: Jev-Omni (akhilaaa3, Gemma-4-12B merged). Published attribution: akhilaaa3.

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.375 s; p95 0.908 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 71.30; C 71.50.

choice: open 83.37; sealed 78.70 (chance-corrected competence).

noul: open 55.28; sealed 56.73 (chance-corrected competence).

score: open 72.79; sealed 76.07 (chance-corrected competence).

71.5095% CI 70.21–72.4070.4982.6084.6856.0570.500.02917estimate
7decider-4b v2Mapika · system-one-open
Evaluation setup

Source entry: decider-4b v2 (Mapika). Published attribution: Mapika.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.203 s; p95 0.405 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 67.53; C 61.59.

choice: open 80.43; sealed 77.31 (chance-corrected competence).

noul: open 42.92; sealed 18.73 (chance-corrected competence).

score: open 61.88; sealed 53.36 (chance-corrected competence).

71.2895% CI 69.11–72.3555.7785.5990.8664.5449.800.01520estimate
8Decision 4B v1.2FlyMy.AI · unclassified
Evaluation setup

Source entry: Decision 4B v1.2 (FlyMyJev, Qwen3.5-4B + LoRA). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.183 s; p95 0.244 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 66.57; C 56.66.

choice: open 80.87; sealed 73.76 (chance-corrected competence).

noul: open 20.13; sealed 16.44 (chance-corrected competence).

score: open 67.79; sealed 63.00 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

70.8395% CI 68.52–71.9753.6688.5693.5063.0751.060.01702estimate
9Imajev-4B (RTX 5090)mohit67890 · unclassified
Evaluation setup

Source entry: Imajev-4B (RTX 5090). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (RTX 5090), offline read-only container

Adjusted latency: p50 0.234 s; p95 0.329 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE (I-2): measured input tokens; zero generated output tokens for signed logits readout. M2 floor uses the 25 Sep 2026 DeepInfra Qwen3.5-4B snapshot rates (USD 0.03/M input, USD 0.15/M output); the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen3.5-9B.

Secondary composites: B 66.20; C 55.91.

choice: open 80.12; sealed 82.33 (chance-corrected competence).

noul: open 21.82; sealed 7.56 (chance-corrected competence).

score: open 66.80; sealed 62.21 (chance-corrected competence).

v1.5 roster addendum A2. No pairwise comparison is inferred for this addition.

70.3995% CI 67.80–71.6153.4788.1291.1363.2650.700.01677estimate
10Decision 4B v1.1FlyMy.AI · unclassified
Evaluation setup

Source entry: Decision 4B v1.1 (FlyMyJev, Qwen3.5-4B + LoRA). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.183 s; p95 0.243 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 66.09; C 55.14.

choice: open 80.12; sealed 74.77 (chance-corrected competence).

noul: open 18.02; sealed 15.75 (chance-corrected competence).

score: open 65.82; sealed 64.17 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

70.3995% CI 66.86–71.5753.1187.3193.5263.0751.560.01702estimate
11Manchego v2.1oraculumai · system-one-open
Evaluation setup

Source entry: Manchego v2.1. Published attribution: oraculumai.

Model author or lab

evaluator-owned offline RTX6000 GPU

Adjusted latency: p50 0.319 s; p95 0.430 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: frozen 25 Sep DeepInfra Qwen/Qwen3.5-4B market reference, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B). $0.03/M input, zero generated output; 836,500 measured input tokens across 1,624 decisions. Frozen v1.5 base-model floor applied; no bookable Manchego tariff claimed. Re-score if the basis changes.

Secondary composites: B 64.38; C 50.11.

choice: open 66.94; sealed 65.04 (chance-corrected competence).

noul: open 34.26; sealed 16.84 (chance-corrected competence).

score: open 62.03; sealed 62.10 (chance-corrected competence).

v1.5 roster addendum A5. No pairwise comparison is inferred for this addition.

68.8195% CI 59.64–70.2851.2084.9388.6364.3347.990.01545estimate
12 ≈SemIf Qwen3.5-4BTheodore Lee (TheoLeeCJ) · jev-rebuild
Evaluation setup

Source entry: SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ). Published attribution: Theodore Lee (TheoLeeCJ).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.229 s; p95 0.356 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 64.31; C 50.22.

choice: open 73.80; sealed 70.96 (chance-corrected competence).

noul: open 30.30; sealed 13.93 (chance-corrected competence).

score: open 57.89; sealed 61.00 (chance-corrected competence).

68.6695% CI 60.06–70.1651.3183.9790.8963.0748.630.01702estimate
13 ≈spark-s1-4b-v6Abhishek Rai (abhishek085) · jev-rebuild
Evaluation setup

Source entry: spark-s1-4b-v6 (Open Spark Jev, abhishek085). Published attribution: Abhishek Rai (abhishek085).

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.406 s; p95 0.690 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 66.86; C 68.16.

choice: open 72.55; sealed 73.11 (chance-corrected competence).

noul: open 50.44; sealed 46.47 (chance-corrected competence).

score: open 64.39; sealed 65.74 (chance-corrected competence).

68.1695% CI 66.27–69.7062.1269.7185.5360.4261.770.02086estimate
14 ≈metask-jev-4bWayfind (metask-ai) · jev-rebuild
Evaluation setup

Source entry: metask-jev-4b. Published attribution: Wayfind (metask-ai).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.285 s; p95 0.394 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 64.12; C 53.61.

choice: open 70.11; sealed 72.41 (chance-corrected competence).

noul: open 43.23; sealed 25.20 (chance-corrected competence).

score: open 56.88; sealed 53.04 (chance-corrected competence).

67.4895% CI 65.39–68.6953.4882.7489.4957.7350.220.02564estimate
15 ≈HopperHopitAI · jev-rebuild
Evaluation setup

Source entry: Hopper. Published attribution: HopitAI.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.395 s; p95 0.479 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 62.93; C 46.85.

choice: open 70.12; sealed 75.48 (chance-corrected competence).

noul: open 10.23; sealed 12.76 (chance-corrected competence).

score: open 65.76; sealed 64.82 (chance-corrected competence).

67.4795% CI 56.34–69.1549.8687.8787.2362.2651.020.01812estimate
16Malkuth-4Bnewfull5 (dhtocks) · jev-rebuild
Evaluation setup

Source entry: Malkuth-4B (newfull5, Kev post-train). Published attribution: newfull5 (dhtocks).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.442 s; p95 0.548 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 63.88; C 54.99.

choice: open 71.17; sealed 69.55 (chance-corrected competence).

noul: open 39.48; sealed 22.84 (chance-corrected competence).

score: open 64.44; sealed 59.23 (chance-corrected competence).

66.7795% CI 64.94–67.8954.4583.2186.1655.8250.540.02968estimate
17Surogate Rune 26B-A4B v3 (RTX PRO 6000)Surogate · unclassified
Evaluation setup

Source entry: Surogate Rune 26B-A4B v3 (RTX PRO 6000). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (RTX PRO 6000), offline read-only container

Adjusted latency: p50 0.351 s; p95 0.721 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.

Secondary composites: B 66.55; C 66.47.

choice: open 81.99; sealed 82.17 (chance-corrected competence).

noul: open 45.59; sealed 49.53 (chance-corrected competence).

score: open 79.28; sealed 79.88 (chance-corrected competence).

v1.5 roster addendum A2. No pairwise comparison is inferred for this addition.

66.4795% CI 65.43–67.0169.7488.3085.9748.9770.530.05024estimate
18 ≈reflex 4Bkshetrajna12 · jev-rebuild
Evaluation setup

Source entry: reflex 4B (kshetrajna12). Published attribution: kshetrajna12.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 2.865 s; p95 4.627 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 61.93; C 48.05.

choice: open 73.66; sealed 69.31 (chance-corrected competence).

noul: open 33.33; sealed 19.45 (chance-corrected competence).

score: open 58.62; sealed 54.56 (chance-corrected competence).

65.2495% CI 58.56–66.3551.4986.8468.7763.1447.780.01693estimate
19 ≈jev-local Qwen3.5-9Bus (GitHub) · jev-rebuild
Evaluation setup

Source entry: jev-local (Qwen3.5-9B). Published attribution: us (GitHub).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.944 s; p95 5.339 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 63.22; C 57.39.

choice: open 66.88; sealed 65.39 (chance-corrected competence).

noul: open 46.94; sealed 37.02 (chance-corrected competence).

score: open 60.71; sealed 60.71 (chance-corrected competence).

65.2495% CI 63.46–66.5056.2877.8272.9858.8654.370.02352estimate
20 ≈djevMaisa · jev-rebuild
Evaluation setup

Source entry: djev (Maisa, diffusion-gemma). Published attribution: Maisa (David Villalón).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.251 s; p95 0.317 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 64.76; C 64.16.

choice: open 76.62; sealed 72.78 (chance-corrected competence).

noul: open 66.16; sealed 70.25 (chance-corrected competence).

score: open 75.33; sealed 72.80 (chance-corrected competence).

64.1695% CI 63.07–65.0072.3380.4191.0048.2271.950.05321estimate
21 ≈Raw Qwen3 4B Instruct 2507 direct logitsAlibaba · raw-logit-control
Evaluation setup

Source entry: Raw Qwen3 4B Instruct 2507 direct logits. Published attribution: Alibaba Qwen / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.262 s; p95 0.456 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 60.33; C 50.45.

choice: open 65.50; sealed 63.46 (chance-corrected competence).

noul: open 49.81; sealed 38.04 (chance-corrected competence).

score: open 56.98; sealed 50.60 (chance-corrected competence).

62.1495% CI 59.35–64.2254.0652.9589.2363.3650.700.01665estimate
22 ≈jqv Qwen3-32BOctalab · jev-rebuild
Evaluation setup

Source entry: jqv (Qwen3-32B zero-shot). Published attribution: hjmurmur (Octalab).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.307 s; p95 1.503 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 57.50; C 42.20.

choice: open 73.00; sealed 73.09 (chance-corrected competence).

noul: open 18.67; sealed 10.18 (chance-corrected competence).

score: open 62.35; sealed 57.23 (chance-corrected competence).

60.7795% CI 51.25–64.2649.0986.7783.3651.1746.830.04242estimate
23 ≈JevK5 v0.2.0allebee · jev-rebuild
Evaluation setup

Source entry: JevK5 v0.2.0. Published attribution: allebee.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.216 s; p95 0.379 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 53.55; C 40.36.

choice: open 77.61; sealed 71.76 (chance-corrected competence).

noul: open 12.19; sealed -3.89 (chance-corrected competence).

score: open 60.23; sealed 62.33 (chance-corrected competence).

58.1195% CI 47.56–68.2446.7084.8790.8663.0743.400.01702estimate
24Qwen3.5-9B Jev-like data-mix v2jsaurabh · jev-rebuild
Evaluation setup

Source entry: Qwen3.5-9B Jev-like data-mix v2. Published attribution: jsaurabh.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.581 s; p95 1.087 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 52.49; C 53.02.

choice: open 72.91; sealed 74.77 (chance-corrected competence).

noul: open 42.37; sealed 31.67 (chance-corrected competence).

score: open 72.56; sealed 68.41 (chance-corrected competence).

53.0295% CI 51.88–53.8860.4580.8682.0045.6958.280.06461estimate
25 ≈Standard One 8BStandard Thinking · jev-rebuild
Evaluation setup

Source entry: Standard One 8B (Standard Thinking). Published attribution: Standard Thinking (myeongho12).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.203 s; p95 0.278 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 47.13; C 47.15.

choice: open 73.72; sealed 72.18 (chance-corrected competence).

noul: open 45.04; sealed 42.98 (chance-corrected competence).

score: open 61.11; sealed 62.56 (chance-corrected competence).

47.7995% CI 46.67–48.5459.6083.1692.4843.2859.240.07772estimate
26NInfer Qwen3.8-Flash-Next mixedIgor L. / NInfer contributors · native-logit
Evaluation setup

Source entry: NInfer Qwen3.8-Flash-Next mixed. Published attribution: Igor L. / NInfer contributors.

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.310 s; p95 0.445 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 47.70; C 47.47.

choice: open 84.93; sealed 79.11 (chance-corrected competence).

noul: open 51.40; sealed 43.02 (chance-corrected competence).

score: open 74.44; sealed 70.30 (chance-corrected competence).

47.4795% CI 46.70–47.8867.2088.5488.6142.5364.140.08233estimate
27Instinct Dual 4BZooWork · decision-apiAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Instinct Dual 4B. Published attribution: ZooWork / pierre-srp.

Model author or lab

operator-hosted free-preview API

Adjusted latency: p50 0.506 s; p95 1.260 s. x2 (demo assumption, not measured)

ESTIMATE: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule); reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B.

Secondary composites: B 42.97; C 32.62.

choice: open 75.05; sealed 74.49 (chance-corrected competence).

noul: open 4.41; sealed -12.80 (chance-corrected competence).

score: open 59.41; sealed 58.15 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

46.9795% CI 37.86–56.7743.1288.3081.9560.1939.950.02124estimate
28 ≈swanOneblockbrain · system-one-open
Evaluation setup

Source entry: swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4). Published attribution: blockbrain.

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.568 s; p95 0.591 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 47.32; C 46.56.

choice: open 88.11; sealed 82.65 (chance-corrected competence).

noul: open 45.55; sealed 59.89 (chance-corrected competence).

score: open 77.84; sealed 73.31 (chance-corrected competence).

46.5695% CI 45.81–46.9471.2287.0684.7442.1571.950.08478estimate
29 ≈Raw Qwen3 8B direct logitsAlibaba · raw-logit-control
Evaluation setup

Source entry: Raw Qwen3 8B direct logits. Published attribution: Alibaba Qwen / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.333 s; p95 0.621 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 44.63; C 32.85.

choice: open 64.86; sealed 56.55 (chance-corrected competence).

noul: open 48.07; sealed 38.40 (chance-corrected competence).

score: open 57.40; sealed 41.56 (chance-corrected competence).

45.2295% CI 36.89–46.7451.1449.1986.8445.5445.500.06539estimate
30 ≈decider-2bMapika · jev-rebuild
Evaluation setup

Source entry: decider-2b (Mapika). Published attribution: Mapika.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.177 s; p95 0.205 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 41.11; C 31.32.

choice: open 52.02; sealed 53.48 (chance-corrected competence).

noul: open 35.24; sealed 5.78 (chance-corrected competence).

score: open 59.11; sealed 48.39 (chance-corrected competence).

45.1095% CI 33.82–54.2642.3471.5394.4064.9235.890.01477estimate
31 ≈system-one Qwen3-8BSean Goedecke · jev-rebuild
Evaluation setup

Source entry: system-one (Qwen3-8B, Sean Goedecke). Published attribution: Sean Goedecke.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.227 s; p95 0.360 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 43.43; C 31.23.

choice: open 63.18; sealed 56.68 (chance-corrected competence).

noul: open 49.89; sealed 40.04 (chance-corrected competence).

score: open 47.55; sealed 45.51 (chance-corrected competence).

44.1495% CI 36.79–45.6650.4749.3890.8844.9747.410.06829estimate
32 ≈system-one-openmithalouni · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: system-one-open (Gemma 4 E2B LoRA on an L4). Published attribution: mithalouni.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 1.150 s; p95 1.295 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 38.80; C 29.47.

choice: open 59.92; sealed 52.04 (chance-corrected competence).

noul: open 15.40; sealed 14.98 (chance-corrected competence).

score: open 57.99; sealed 49.52 (chance-corrected competence).

42.4495% CI 33.84–51.9241.6471.6378.2768.3338.850.01137estimate
33 ≈Autoloops Gemma 4 31B ITAutoloops · jevAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Autoloops – Gemma 4 31B IT. Published attribution: Autoloops.

the operator's hosted API

Adjusted latency: p50 0.607 s; p95 0.667 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1)

Secondary composites: B 41.84; C 40.53.

choice: open 86.36; sealed 87.51 (chance-corrected competence).

noul: open 68.73; sealed 67.45 (chance-corrected competence).

score: open 75.55; sealed 74.51 (chance-corrected competence).

40.5395% CI 40.09–40.7676.6885.8483.9239.5976.490.10323tariff
34GPT-6 Luna (low)OpenAI · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: GPT-6 Luna (low reasoning effort). Published attribution: OpenAI.

the operator's hosted API

Adjusted latency: p50 1.580 s; p95 3.031 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1); no exact base-model floor applies

Secondary composites: B 43.10; C 40.48.

choice: open 99.06; sealed 97.87 (chance-corrected competence).

noul: open 85.25; sealed 90.18 (chance-corrected competence).

score: open 99.87; sealed 99.52 (chance-corrected competence).

40.4895% CI 40.27–40.6695.2994.9273.2039.0695.860.10750estimate
35 ≈GPT-6 Luna (medium)OpenAI · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: GPT-6 Luna (default medium reasoning effort). Published attribution: OpenAI.

the operator's hosted API

Adjusted latency: p50 1.555 s; p95 3.088 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1); no exact base-model floor applies

Secondary composites: B 41.35; C 38.75.

choice: open 99.38; sealed 99.51 (chance-corrected competence).

noul: open 90.69; sealed 89.93 (chance-corrected competence).

score: open 98.89; sealed 99.02 (chance-corrected competence).

38.7595% CI 38.57–38.9396.2495.5673.1938.3296.150.11379estimate
36 ≈JevOneJuspay · jev-rebuild
Evaluation setup

Source entry: JevOne. Published attribution: Juspay.

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.267 s; p95 0.339 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 37.40; C 31.13.

choice: open 82.43; sealed 77.14 (chance-corrected competence).

noul: open 9.38; sealed 17.13 (chance-corrected competence).

score: open 69.91; sealed 68.81 (chance-corrected competence).

38.2595% CI 37.38–38.7954.1384.9490.4339.8454.360.10122estimate
37 ≈kev 4BJared Palmer · jev-rebuild
Evaluation setup

Source entry: kev 4B (research preview). Published attribution: Jared Palmer.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.493 s; p95 0.574 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 34.59; C 26.44.

choice: open 53.89; sealed 46.76 (chance-corrected competence).

noul: open 29.06; sealed 4.29 (chance-corrected competence).

score: open 58.24; sealed 48.55 (chance-corrected competence).

38.0795% CI 27.48–46.7739.8667.6485.4865.7833.200.01382estimate
38 ≈kev 8BJared Palmer · jev-rebuild
Evaluation setup

Source entry: kev 8B (research preview). Published attribution: Jared Palmer.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.511 s; p95 0.754 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 33.09; C 23.72.

choice: open 60.67; sealed 58.72 (chance-corrected competence).

noul: open 38.20; sealed 14.00 (chance-corrected competence).

score: open 64.04; sealed 54.14 (chance-corrected competence).

34.1595% CI 26.79–37.4848.3071.0284.1440.4242.280.09682estimate
39 ≈open-alternative-jev Qwen3.5-4BIkerMoel · jev-rebuild
Evaluation setup

Source entry: open-alternative-jev (Qwen3.5-4B, IkerMoel). Published attribution: IkerMoel.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.223 s; p95 0.334 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE (proxy tokens, I-2): exact base-model market reference; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 29.92; C 23.31.

choice: open 63.28; sealed 65.00 (chance-corrected competence).

noul: open 5.29; sealed -16.76 (chance-corrected competence).

score: open 59.22; sealed 48.12 (chance-corrected competence).

33.5695% CI 25.75–41.6637.3676.8791.2963.2832.120.01675estimate
40 ≈Bespoke Nimble 9BBespoke Labs · jev-rebuild
Evaluation setup

Source entry: Bespoke Nimble 9B (Bespoke Labs). Published attribution: Bespoke Labs.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.547 s; p95 0.927 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 32.31; C 31.82.

choice: open 71.68; sealed 72.82 (chance-corrected competence).

noul: open 52.86; sealed 44.55 (chance-corrected competence).

score: open 70.90; sealed 69.68 (chance-corrected competence).

31.8295% CI 31.19–32.2763.7577.1682.9536.7562.350.12833estimate
41 ≈Malkuth-2Bnewfull5 (dhtocks) · jev-rebuild
Evaluation setup

Source entry: Malkuth-2B (newfull5, Kev post-train). Published attribution: newfull5 (dhtocks).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.231 s; p95 0.286 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 26.38; C 20.76.

choice: open 54.83; sealed 48.49 (chance-corrected competence).

noul: open 26.01; sealed -5.78 (chance-corrected competence).

score: open 51.90; sealed 43.12 (chance-corrected competence).

29.9095% CI 20.96–38.9835.5475.1991.8065.5728.610.01405estimate
42openjev-sglang Qwen3.6-35B-A3Bekzhang · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: openjev-sglang (Qwen3.6-35B-A3B on SGLang). Published attribution: ekzhang.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 1.196 s; p95 1.305 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 29.18; C 27.73.

choice: open 80.55; sealed 74.60 (chance-corrected competence).

noul: open 36.38; sealed 27.20 (chance-corrected competence).

score: open 67.14; sealed 66.02 (chance-corrected competence).

29.0395% CI 28.42–29.4058.6582.8778.0735.6355.940.13981estimate
43 ≈decider-35b-a3bMapika · jev-rebuild
Evaluation setup

Source entry: decider-35b-a3b (Mapika). Published attribution: Mapika.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.241 s; p95 0.330 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 27.71; C 27.50.

choice: open 74.30; sealed 76.47 (chance-corrected competence).

noul: open 51.34; sealed 31.31 (chance-corrected competence).

score: open 66.79; sealed 62.55 (chance-corrected competence).

27.5095% CI 26.93–27.8760.4682.0491.0034.3956.780.15387estimate
44local-jev Qwen3.5-4BAmith Chandrappa (amithgc) · jev-rebuild
Evaluation setup

Source entry: local-jev Qwen3.5-4B. Published attribution: Amith Chandrappa (amithgc).

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.427 s; p95 0.991 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 22.73; C 17.93.

choice: open 73.67; sealed 69.06 (chance-corrected competence).

noul: open -23.24; sealed -38.29 (chance-corrected competence).

score: open 59.68; sealed 61.61 (chance-corrected competence).

25.8295% CI 19.89–32.9633.7582.4583.7459.2730.790.02279estimate
45Nemotron Diffusion 8B (optimized vLLM)pst2154 · system-one-open
Evaluation setup

Source entry: Nemotron Diffusion 8B (pst2154, optimized vLLM). Published attribution: pst2154.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.187 s; p95 0.225 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 22.70; C 17.84.

choice: open 57.05; sealed 58.81 (chance-corrected competence).

noul: open 10.77; sealed -16.91 (chance-corrected competence).

score: open 50.79; sealed 42.55 (chance-corrected competence).

v1.5 roster addendum A3. No pairwise comparison is inferred for this addition.

25.6895% CI 18.60–32.7833.8476.2593.7455.4928.150.03045estimate
46 ≈Open-Jev 9BZefan Cai (@Zefan_Cai) · jev-rebuild
Evaluation setup

Source entry: Open-Jev 9B (Zefan Cai). Published attribution: Zefan Cai (@Zefan_Cai).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 1.276 s; p95 3.432 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 24.99; C 24.36.

choice: open 66.19; sealed 71.45 (chance-corrected competence).

noul: open 48.96; sealed 58.69 (chance-corrected competence).

score: open 70.21; sealed 67.42 (chance-corrected competence).

24.3695% CI 23.93–24.6463.8281.5373.5933.0565.860.17041estimate
47 ≈Decision 2B v59FlyMy.AI · jev-rebuild
Evaluation setup

Source entry: Decision 2B (FlyMy.AI, v59). Published attribution: FlyMy.AI (@denti).

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.300 s; p95 0.307 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 19.30; C 15.64.

choice: open 58.27; sealed 67.70 (chance-corrected competence).

noul: open -16.11; sealed -35.13 (chance-corrected competence).

score: open 61.78; sealed 51.36 (chance-corrected competence).

22.5295% CI 16.68–28.7731.3186.2290.3666.3727.980.01321estimate
48 ≈GPT-5.6 Luna (low)OpenAI · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: GPT-5.6 Luna (low reasoning effort). Published attribution: OpenAI.

the operator's hosted API

Adjusted latency: p50 1.321 s; p95 2.970 s. none (hosted API, measured as is)

operator list price

Secondary composites: B 24.16; C 22.37.

choice: open 95.94; sealed 93.85 (chance-corrected competence).

noul: open 87.11; sealed 94.15 (chance-corrected competence).

score: open 96.42; sealed 98.52 (chance-corrected competence).

22.3795% CI 22.26–22.4694.3394.7174.0630.6795.510.20467tariff
49 ≈typecastlmMikhail Gribov · system-one-open
Evaluation setup

Source entry: typecastlm (Mikhail Gribov, Qwen3.5-4B computed head). Published attribution: Mikhail Gribov.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.225 s; p95 0.278 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 18.85; C 15.16.

choice: open 66.45; sealed 72.01 (chance-corrected competence).

noul: open -52.09; sealed -7.49 (chance-corrected competence).

score: open 56.73; sealed 51.83 (chance-corrected competence).

21.8395% CI 16.23–28.5831.2476.7192.0463.9638.780.01590estimate
50 ≈JEV Qwen3.5-9B Base NVFP4WilfLin · jev-rebuild
Evaluation setup

Source entry: JEV Qwen3.5-9B Base NVFP4. Published attribution: WilfLin.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.190 s; p95 0.234 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE (proxy tokens, I-2): exact base-model market reference

Secondary composites: B 17.76; C 13.94.

choice: open 71.17; sealed 59.87 (chance-corrected competence).

noul: open -34.85; sealed -22.47 (chance-corrected competence).

score: open 62.32; sealed 57.43 (chance-corrected competence).

20.0795% CI 15.19–25.5132.2480.9393.5447.5931.610.05584estimate
51Gemini 3.1 Flash-LiteGoogle · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Gemini 3.1 Flash-Lite. Published attribution: Google.

the operator's hosted API

Adjusted latency: p50 0.863 s; p95 1.144 s. none (hosted API, measured as is)

operator list price

Secondary composites: B 20.78; C 19.58.

choice: open 83.98; sealed 84.09 (chance-corrected competence).

noul: open 73.61; sealed 75.05 (chance-corrected competence).

score: open 72.25; sealed 76.53 (chance-corrected competence).

19.5895% CI 19.30–19.8377.5974.6880.0529.7678.560.21941tariff
52AutoJev-27Bdenis-pplx · unclassified
Evaluation setup

Source entry: AutoJev-27B (denis-pplx, Qwen3.8-27B). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.352 s; p95 0.492 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 20.45; C 19.54.

choice: open 85.24; sealed 82.24 (chance-corrected competence).

noul: open 50.44; sealed 62.00 (chance-corrected competence).

score: open 80.06; sealed 76.70 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

19.5495% CI 19.30–19.6372.7887.7087.6229.3773.650.22619estimate
53AutoJev-27B (RTX PRO 6000)denis-pplx · unclassified
Evaluation setup

Source entry: AutoJev-27B (RTX PRO 6000). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (RTX PRO 6000), offline read-only container

Adjusted latency: p50 0.326 s; p95 0.602 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.

Secondary composites: B 20.39; C 19.48.

choice: open 84.93; sealed 81.70 (chance-corrected competence).

noul: open 51.14; sealed 62.00 (chance-corrected competence).

score: open 79.98; sealed 76.76 (chance-corrected competence).

v1.5 roster addendum A2. No pairwise comparison is inferred for this addition.

19.4895% CI 19.26–19.6072.7586.6887.0729.3773.490.22619estimate
54NInfer Qwen3.8-27B NVFP4Igor L. / NInfer contributors · native-logit
Evaluation setup

Source entry: NInfer Qwen3.8-27B NVFP4. Published attribution: Igor L. / NInfer contributors.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.248 s; p95 0.416 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 19.35; C 18.74.

choice: open 80.86; sealed 80.00 (chance-corrected competence).

noul: open 47.77; sealed 44.73 (chance-corrected competence).

score: open 72.29; sealed 67.19 (chance-corrected competence).

18.7495% CI 18.45–18.9265.4785.9189.8629.1263.970.23052estimate
55Eikos-27Bcaiovicentino1 · unclassified
Evaluation setup

Source entry: Eikos-27B (caiovicentino1, Qwen3.8-27B). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.354 s; p95 0.488 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 19.48; C 18.50.

choice: open 87.62; sealed 88.27 (chance-corrected competence).

noul: open 56.47; sealed 64.15 (chance-corrected competence).

score: open 78.35; sealed 75.66 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

18.5095% CI 18.28–18.6075.0886.2987.6328.6976.020.23824estimate
56 ≈NInfer Qwen3.8-27B NVFP4 (T=1.5)Igor L. / NInfer contributors · native-logit
Evaluation setup

Source entry: NInfer Qwen3.8-27B NVFP4 (T=1.5). Published attribution: Igor L. / NInfer contributors.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.248 s; p95 0.416 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 18.89; C 18.48.

choice: open 80.86; sealed 80.00 (chance-corrected competence).

noul: open 36.83; sealed 36.36 (chance-corrected competence).

score: open 69.61; sealed 62.80 (chance-corrected competence).

18.4895% CI 18.18–18.6561.0886.4989.8629.1259.720.23052estimate
57Instinct Qwen3.8-27BZooWork · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Instinct (ZooWork, Qwen3.8-27B). Published attribution: rayrain-srp (ZooWork).

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 0.519 s; p95 1.308 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 18.86; C 18.32.

choice: open 84.49; sealed 79.67 (chance-corrected competence).

noul: open 28.50; sealed 39.56 (chance-corrected competence).

score: open 72.75; sealed 71.52 (chance-corrected competence).

18.3295% CI 18.03–18.5162.7584.9681.6829.1663.580.22979estimate
58OpenJev (thinking, BF16)razorback16 · jev-rebuild
Evaluation setup

Source entry: OpenJev (thinking, BF16). Published attribution: razorback16.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 1.561 s; p95 2.632 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 19.27; C 17.94.

choice: open 81.67; sealed 78.51 (chance-corrected competence).

noul: open 80.61; sealed 88.62 (chance-corrected competence).

score: open 85.26; sealed 90.57 (chance-corrected competence).

17.9495% CI 17.76–18.0884.2183.1073.8628.5185.900.24146estimate
59djev (thinking)Maisa · jev-rebuild
Evaluation setup

Source entry: djev (thinking). Published attribution: David Villalon / Maisa.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 1.481 s; p95 3.959 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 18.47; C 17.40.

choice: open 64.77; sealed 48.29 (chance-corrected competence).

noul: open 74.63; sealed 95.75 (chance-corrected competence).

score: open 90.27; sealed 90.39 (chance-corrected competence).

17.4095% CI 17.28–17.5177.3595.6672.3228.1378.140.24866estimate
60LitJev Qwen3.8-27BZhengxu Yu · jev-rebuild
Evaluation setup

Source entry: LitJev (Qwen3.8-27B). Published attribution: Zhengxu Yu.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 3.068 s; p95 4.906 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 16.74; C 15.41.

choice: open 80.44; sealed 79.03 (chance-corrected competence).

noul: open 23.49; sealed 33.71 (chance-corrected competence).

score: open 70.26; sealed 63.02 (chance-corrected competence).

16.3195% CI 16.03–16.4858.3284.4568.2228.3658.590.24438estimate
61Bev / Bonsai 27BReza Sayar · system-one-open
Evaluation setup

Source entry: Bev / Bonsai 27B. Published attribution: Reza Sayar.

Model author or lab

evaluator-owned H100 GPU

Adjusted latency: p50 2.191 s; p95 2.433 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate. Frozen market reference qwen/qwen3.8-27b.

Secondary composites: B 15.99; C 12.34.

choice: open 74.11; sealed 76.27 (chance-corrected competence).

noul: open 4.29; sealed 29.49 (chance-corrected competence).

score: open 67.71; sealed 66.61 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

15.7795% CI 15.13–16.0553.0877.6872.7328.2357.460.24676estimate
62 ≈Raw Phi-4 mini direct logitsMicrosoft · raw-logit-control
Evaluation setup

Source entry: Raw Phi-4 mini direct logits. Published attribution: Microsoft / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.281 s; p95 0.431 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 13.11; C 10.57.

choice: open 54.10; sealed 56.67 (chance-corrected competence).

noul: open 6.03; sealed -43.64 (chance-corrected competence).

score: open 54.69; sealed 46.85 (chance-corrected competence).

15.2295% CI 10.26–20.7427.6271.4189.1753.2219.960.03625estimate
63 ≈OpenSourceJev Qwen3.5-4B Q4_K_Msabeel111 · jev-rebuild
Evaluation setup

Source entry: OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp). Published attribution: sabeel111.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 1.389 s; p95 4.041 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 11.16; C 9.20.

choice: open 64.38; sealed 51.49 (chance-corrected competence).

noul: open -36.12; sealed -37.24 (chance-corrected competence).

score: open 55.31; sealed 56.79 (chance-corrected competence).

13.2595% CI 8.98–18.3125.7776.0472.5169.1423.680.01069estimate
64reflex-27bkshetrajna12 · jev-rebuild
Evaluation setup

Source entry: reflex-27b (Qwen3.8-27B). Published attribution: kshetrajna12.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 2.748 s; p95 4.379 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 13.77; C 13.19.

choice: open 84.80; sealed 80.30 (chance-corrected competence).

noul: open 30.92; sealed 35.71 (chance-corrected competence).

score: open 73.55; sealed 71.49 (chance-corrected competence).

13.1995% CI 12.99–13.3062.7985.9369.2025.8062.500.29734estimate
65 ≈Open-Jev 2BZefan Cai (@Zefan_Cai) · jev-rebuild
Evaluation setup

Source entry: Open-Jev 2B (Zefan Cai). Published attribution: Zefan Cai (@Zefan_Cai).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 1.008 s; p95 2.631 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 8.44; C 6.30.

choice: open 51.52; sealed 45.49 (chance-corrected competence).

noul: open 17.34; sealed -8.15 (chance-corrected competence).

score: open 47.50; sealed 47.69 (chance-corrected competence).

9.0795% CI 6.62–11.5433.5673.6275.7633.0528.340.17041estimate
66 ≈GLiNER2 largeFastino · classifier
Evaluation setup

Source entry: GLiNER2 large (Fastino). Published attribution: Fastino.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 1.879 s; p95 17.719 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 7.19; C 5.83.

choice: open 26.57; sealed 28.42 (chance-corrected competence).

noul: open 16.35; sealed -0.62 (chance-corrected competence).

score: open 35.16; sealed 29.07 (chance-corrected competence).

8.4095% CI 5.21–12.3422.4942.3864.7877.5918.960.00558estimate
67 ≈Qwen3-Reranker-4BAlibaba · reranker
Evaluation setup

Source entry: Qwen3-Reranker-4B. Published attribution: Qwen.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.511 s; p95 2.140 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: hosted exact-model reference (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies

Secondary composites: B 5.86; C 4.90.

choice: open 56.70; sealed 45.60 (chance-corrected competence).

noul: open -10.06; sealed -45.93 (chance-corrected competence).

score: open 44.40; sealed 40.42 (chance-corrected competence).

7.0695% CI 4.26–10.7421.0376.2679.6148.4013.360.05249estimate
68 ≈DeepSeek V4.1 FlashDeepSeek · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: DeepSeek V4.1 Flash (thinking default). Published attribution: DeepSeek.

the operator's hosted API

Adjusted latency: p50 1.776 s; p95 6.465 s. none (hosted API, measured as is)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 7.41; C 6.65.

choice: open 96.20; sealed 98.23 (chance-corrected competence).

noul: open 86.17; sealed 96.55 (chance-corrected competence).

score: open 93.16; sealed 91.82 (chance-corrected competence).

6.6595% CI 6.62–6.6693.6996.9269.4019.0995.530.49760estimate
69 ≈SimpleJev Qwen3.5-0.8BFeatherless AI · jev-rebuild
Evaluation setup

Source entry: SimpleJev (Qwen3.5-0.8B, CPU). Published attribution: sabeel111 / Featherless AI.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 7.371 s; p95 16.793 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 3.26; C 2.77.

choice: open 28.28; sealed 27.77 (chance-corrected competence).

noul: open -12.83; sealed 4.25 (chance-corrected competence).

score: open 24.54; sealed 28.47 (chance-corrected competence).

4.0095% CI 2.06–6.6616.7546.6759.0770.3120.160.00976estimate
70 ≈SimpleJev Qwen3.8-27BFeatherless AI · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: SimpleJev Qwen3.8-27B. Published attribution: Featherless AI.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 1.679 s; p95 1.920 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 3.71; C 3.36.

choice: open 85.43; sealed 80.67 (chance-corrected competence).

noul: open 51.10; sealed 60.95 (chance-corrected competence).

score: open 79.20; sealed 79.18 (chance-corrected competence).

3.3695% CI 3.33–3.3772.7587.2074.9214.9073.600.68680estimate
71 ≈decision-machine-1milliseconds.ai · decision-apiAPI: sealed item text sent to the operator
Evaluation setup

Source entry: decision-machine-1 (milliseconds.ai). Published attribution: milliseconds.ai (Baptiste Laget).

Model author or lab

the operator's hosted API

Adjusted latency: p50 0.180 s; p95 0.293 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1); no exact base-model floor applies

Secondary composites: B 2.50; C 2.25.

choice: open 53.27; sealed 44.94 (chance-corrected competence).

noul: open -17.55; sealed -67.96 (chance-corrected competence).

score: open 46.78; sealed 38.41 (chance-corrected competence).

3.2495% CI 1.73–5.3914.8281.1892.7756.305.130.02863estimate
72GLiNER2.5 multiFastino · classifier
Evaluation setup

Source entry: GLiNER2.5 multi (Fastino, 287M). Published attribution: Fastino.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 1.370 s; p95 16.069 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 2.10; C 1.89.

choice: open 19.44; sealed 17.24 (chance-corrected competence).

noul: open 0.16; sealed -2.58 (chance-corrected competence).

score: open 27.74; sealed 21.98 (chance-corrected competence).

2.7295% CI 1.20–5.1014.0057.9666.5786.6212.210.00279estimate
73Bosun v3.1 0.6BClause Logic · system-one-open
Evaluation setup

Source entry: Bosun v3.1 0.6B. Published attribution: Clause Logic.

Model author or lab

evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1, local_files_only=True), credential-free (no API key, no HF token)

Adjusted latency: p50 4.037 s; p95 16.515 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-base frozen raw-qwen3-0.6b estimate.

Secondary composites: B 1.92; C 1.74.

choice: open 33.44; sealed 36.69 (chance-corrected competence).

noul: open -4.15; sealed -57.85 (chance-corrected competence).

score: open 40.17; sealed 37.15 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

2.5095% CI 1.15–4.5613.5865.1661.7677.555.330.00560estimate
74GLiNER2.5 baseFastino · classifier
Evaluation setup

Source entry: GLiNER2 (Fastino, gliner2.5-base). Published attribution: Fastino.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.998 s; p95 9.285 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 1.80; C 1.58.

choice: open 24.80; sealed 20.99 (chance-corrected competence).

noul: open 1.50; sealed -10.00 (chance-corrected competence).

score: open 25.42; sealed 18.14 (chance-corrected competence).

2.2795% CI 0.90–4.5013.4835.5770.3386.629.710.00279estimate
75Deem 0.8B v1LibertAI · system-one-open
Evaluation setup

Source entry: Deem 0.8B v1. Published attribution: LibertAI.

Model author or lab

evaluator-owned H100 GPU

Adjusted latency: p50 0.501 s; p95 0.739 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate. Frozen market reference deepinfra:Qwen/Qwen3.5-0.8B.

Secondary composites: B 1.66; C 1.48.

choice: open 28.48; sealed 19.54 (chance-corrected competence).

noul: open 7.67; sealed -28.73 (chance-corrected competence).

score: open 37.90; sealed 19.95 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

2.1495% CI 0.90–4.3413.0238.8884.3280.293.590.00454estimate
76 ≈JevActeinptein · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: JevAct (einptein, jev1-2b-v2). Published attribution: einptein.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 0.802 s; p95 2.836 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 1.14; C 1.06.

choice: open 31.06; sealed 35.28 (chance-corrected competence).

noul: open -10.65; sealed -55.35 (chance-corrected competence).

score: open 38.74; sealed 30.47 (chance-corrected competence).

1.5295% CI 0.52–3.1811.2462.5476.4368.563.470.01117estimate
77 ≈CLM-8BContrastive-LM · system-one-open
Evaluation setup

Source entry: CLM-8B (Contrastive-LM, clm-latest). Published attribution: Contrastive-LM (Kwok, Kang, Suresh, Saad-Falcon, Pavone, Ré, Mirhoseini).

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.183 s; p95 0.256 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 1.16; C 1.05.

choice: open 19.36; sealed 8.29 (chance-corrected competence).

noul: open 6.95; sealed -13.49 (chance-corrected competence).

score: open 24.96; sealed 22.52 (chance-corrected competence).

1.5195% CI 0.48–3.1711.4348.4893.2950.505.770.04466estimate
78 ≈kev 0.6BJared Palmer · jev-rebuild
Evaluation setup

Source entry: kev 0.6B (research preview). Published attribution: Jared Palmer.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.422 s; p95 0.450 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.92; C 0.88.

choice: open 33.25; sealed 21.31 (chance-corrected competence).

noul: open 1.96; sealed -57.60 (chance-corrected competence).

score: open 41.18; sealed 31.81 (chance-corrected competence).

1.2695% CI 0.49–2.5010.3467.6587.2280.09-1.490.00461estimate
79 ≈Raw Qwen3 0.6B direct logitsAlibaba · raw-logit-control
Evaluation setup

Source entry: Raw Qwen3 0.6B direct logits. Published attribution: Alibaba Qwen / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.276 s; p95 0.331 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.86; C 0.75.

choice: open 17.12; sealed 23.80 (chance-corrected competence).

noul: open 13.42; sealed -4.25 (chance-corrected competence).

score: open -0.51; sealed 13.79 (chance-corrected competence).

1.0895% CI 0.27–2.6110.5621.4590.3977.5811.110.00559estimate
80 ≈GLiNER2.5 smallFastino · classifier
Evaluation setup

Source entry: GLiNER2.5 small (Fastino, 74M). Published attribution: Fastino.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.472 s; p95 3.891 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.65; C 0.62.

choice: open 15.71; sealed 11.87 (chance-corrected competence).

noul: open 4.31; sealed -34.11 (chance-corrected competence).

score: open 31.47; sealed 27.19 (chance-corrected competence).

0.8995% CI 0.23–2.209.1955.9077.3686.621.650.00279estimate
81 ≈Raw Qwen3 1.7B direct logitsAlibaba · raw-logit-control
Evaluation setup

Source entry: Raw Qwen3 1.7B direct logits. Published attribution: Alibaba Qwen / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.282 s; p95 0.337 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.67; C 0.59.

choice: open 41.04; sealed 23.29 (chance-corrected competence).

noul: open 11.66; sealed -3.45 (chance-corrected competence).

score: open -15.90; sealed 1.40 (chance-corrected competence).

0.8595% CI 0.15–2.289.6721.5790.2268.557.080.01118estimate
82 ≈MirrorBluusun · jev-rebuild
Evaluation setup

Source entry: Mirror. Published attribution: Bluusun.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 4.578 s; p95 9.436 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.15; C 0.15.

choice: open 1.64; sealed -9.37 (chance-corrected competence).

noul: open -0.61; sealed -4.55 (chance-corrected competence).

score: open 19.21; sealed 27.26 (chance-corrected competence).

0.2295% CI 0.01–0.905.6043.2163.6489.274.450.00228estimate
83 ≈ZeroEntropy zerank-2ZeroEntropy · reranker
Evaluation setup

Source entry: ZeroEntropy zerank-2. Published attribution: ZeroEntropy.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.473 s; p95 1.984 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies

Secondary composites: B 0.09; C 0.09.

choice: open 56.57; sealed 46.96 (chance-corrected competence).

noul: open -68.07; sealed -85.64 (chance-corrected competence).

score: open 41.40; sealed 37.36 (chance-corrected competence).

0.1395% CI 0.01–0.494.7681.7180.2848.40-0.440.05249estimate
84 ≈jeffLogan Markewich · jev-rebuild
Evaluation setup

Source entry: jeff (Logan Markewich, GLiFormer 400M). Published attribution: Logan Markewich.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 7.137 s; p95 37.170 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.08; C 0.08.

choice: open 33.01; sealed 21.79 (chance-corrected competence).

noul: open -28.09; sealed -60.76 (chance-corrected competence).

score: open 33.21; sealed 28.29 (chance-corrected competence).

0.1295% CI 0.00–0.514.4380.2755.7681.09-3.560.00427estimate
85smalljev semantic-v9Aditya (isHeSatoshi) · jev-rebuild
Evaluation setup

Source entry: smalljev semantic-v9. Published attribution: Aditya (isHeSatoshi).

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.451 s; p95 0.501 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.05; C 0.05.

choice: open 31.14; sealed 18.27 (chance-corrected competence).

noul: open -40.40; sealed -63.82 (chance-corrected competence).

score: open 42.06; sealed 34.99 (chance-corrected competence).

0.0795% CI 0.00–0.383.6673.0986.4660.77-3.520.02030estimate
86Laya multilingualConvai Innovations · system-one-open
Evaluation setup

Source entry: Laya multilingual. Published attribution: Convai Innovations.

Model author or lab

evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1)

Adjusted latency: p50 1.036 s; p95 4.097 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor.

Secondary composites: B 0.01; C 0.01.

choice: open 20.06; sealed 9.13 (chance-corrected competence).

noul: open -35.98; sealed -31.05 (chance-corrected competence).

score: open 24.29; sealed 27.68 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

0.0295% CI 0.00–0.272.3543.5873.7282.161.920.00393estimate
87 ≈OpenDecision ModernBERT-largeDeepan Wadhwa · classifier
Evaluation setup

Source entry: OpenDecision (ModernBERT-large zero-shot). Published attribution: Deepan Wadhwa.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.297 s; p95 0.740 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.01; C 0.01.

choice: open 34.17; sealed 20.04 (chance-corrected competence).

noul: open -47.14; sealed -70.91 (chance-corrected competence).

score: open 41.38; sealed 35.08 (chance-corrected competence).

0.0195% CI 0.00–0.192.0772.5286.5879.11-5.260.00497estimate
88 ≈BAAI bge-reranker-v2-m3BAAI · reranker
Evaluation setup

Source entry: BAAI bge-reranker-v2-m3. Published attribution: BAAI.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.209 s; p95 0.422 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: base-model market reference (deepinfra:BAAI/bge-m3)

Secondary composites: B 0.00; C 0.00.

choice: open 4.41; sealed 4.20 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 32.30; sealed 27.38 (chance-corrected competence).

0.0095% CI 0.00–0.000.0083.3390.5559.33-22.800.02269estimate
89 ≈Certo v1AltSlate Labs · jev-rebuild
Evaluation setup

Source entry: Certo v1 (AltSlate Labs). Published attribution: AltSlate Labs.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.264 s; p95 0.273 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open -1.85; sealed -2.05 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 29.49; sealed 27.25 (chance-corrected competence).

0.0095% CI 0.00–0.000.0087.5291.4396.91-24.930.00127estimate
90 ≈Decision Fast v53aFlyMy.AI · jev-rebuild
Evaluation setup

Source entry: Decision Fast (FlyMy.AI, v53a). Published attribution: FlyMy.AI (@denti).

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.266 s; p95 0.270 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 38.36; sealed 18.49 (chance-corrected competence).

noul: open -59.07; sealed -86.84 (chance-corrected competence).

score: open 42.64; sealed 33.26 (chance-corrected competence).

0.0095% CI 0.00–0.000.0076.0891.4480.09-11.700.00461estimate
91 ≈Alibaba GTE Reranker ModernBERT-baseAlibaba · reranker
Evaluation setup

Source entry: Alibaba GTE Reranker ModernBERT-base. Published attribution: Alibaba-NLP.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.212 s; p95 0.336 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: size-class proxy (deepinfra:thenlper/gte-base); no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 13.02; sealed 2.35 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 31.09; sealed 28.18 (chance-corrected competence).

0.0095% CI 0.00–0.000.0075.2091.4969.27-23.160.01058estimate
92 ≈kev 0.5BJared Palmer · jev-rebuild
Evaluation setup

Source entry: kev 0.5B. Published attribution: Jared Palmer.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.357 s; p95 0.400 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 25.24; sealed 11.19 (chance-corrected competence).

noul: open -48.02; sealed -94.55 (chance-corrected competence).

score: open 33.12; sealed 16.56 (chance-corrected competence).

0.0095% CI 0.00–0.000.0064.9888.4580.09-22.260.00461estimate
93 ≈LayaConvai Innovations · jev-rebuild
Evaluation setup

Source entry: Laya (Convai Innovations, ModernBERT-large 421M). Published attribution: Convai Innovations.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 1.485 s; p95 2.747 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 37.10; sealed 12.77 (chance-corrected competence).

noul: open -75.74; sealed -81.53 (chance-corrected competence).

score: open 37.21; sealed 29.20 (chance-corrected competence).

0.0095% CI 0.00–0.000.0073.7173.8984.89-13.190.00319estimate
94 ≈lev-350mFranck Verrot (franckverrot) · jev-rebuild
Evaluation setup

Source entry: lev-350m (Franck Verrot, LFM2.5-350M). Published attribution: Franck Verrot (franckverrot).

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.194 s; p95 0.223 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 33.06; sealed 15.01 (chance-corrected competence).

noul: open -52.20; sealed -100.00 (chance-corrected competence).

score: open 35.72; sealed 30.94 (chance-corrected competence).

0.0095% CI 0.00–0.000.0077.6893.6380.04-18.010.00463estimate
95 ≈Qwen3.5-0.8B Decision ModelMourad Ghafiri · jev-rebuild
Evaluation setup

Source entry: Qwen3.5-0.8B Decision Model (Mourad Ghafiri). Published attribution: Mourad Ghafiri.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 1.308 s; p95 5.023 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 0.00; C 0.00.

choice: open 34.93; sealed 40.44 (chance-corrected competence).

noul: open -69.95; sealed -86.04 (chance-corrected competence).

score: open 45.21; sealed 34.46 (chance-corrected competence).

0.0095% CI 0.00–0.020.0073.5771.8279.52-3.710.00482estimate
96 ≈Mixedbread mxbai-rerank-base-v2Mixedbread · reranker
Evaluation setup

Source entry: Mixedbread mxbai-rerank-base-v2. Published attribution: Mixedbread.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.235 s; p95 0.495 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-0.6B); no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 5.84; sealed 0.62 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 31.19; sealed 27.26 (chance-corrected competence).

0.0095% CI 0.00–0.000.0086.6189.3460.34-24.040.02100estimate
97 ≈Needle 3 (2-bit)Cactus Compute · small-tool-model
Evaluation setup

Source entry: Needle 3 (Cactus, 2-bit, local CPU). Published attribution: Cactus Compute.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 135.513 s; p95 285.427 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open -3.57; sealed -8.14 (chance-corrected competence).

noul: open -14.52; sealed -23.60 (chance-corrected competence).

score: open -15.26; sealed -31.60 (chance-corrected competence).

0.0095% CI 0.00–0.000.000.0034.1361.57-21.110.01910estimate
98 ≈Needle 3 (options as tools)Cactus Compute · small-tool-model
Evaluation setup

Source entry: Needle 3, options as tools (post-hoc adapter mode). Published attribution: Cactus Compute.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 58.160 s; p95 136.690 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 10.87; sealed 2.20 (chance-corrected competence).

noul: open -12.38; sealed -30.11 (chance-corrected competence).

score: open -15.85; sealed -31.79 (chance-corrected competence).

0.0095% CI 0.00–0.000.000.0041.0061.57-19.900.01910estimate
99 ≈open-jev-deberta-v3-largeKotoba Labs · jev-rebuild
Evaluation setup

Source entry: open-jev-deberta-v3-large (local CPU). Published attribution: Kotoba Labs.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 2.933 s; p95 5.097 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 21.00; sealed 16.66 (chance-corrected competence).

noul: open -79.16; sealed -87.38 (chance-corrected competence).

score: open 30.24; sealed 26.89 (chance-corrected competence).

0.0095% CI 0.00–0.000.0077.0668.2577.59-14.610.00558estimate
100 ≈Open Jev JSON CanvasJoshuaSP · jev-rebuild
Evaluation setup

Source entry: Open Jev JSON Canvas (JoshuaSP). Published attribution: JoshuaSP.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.437 s; p95 0.635 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 81.05; sealed 76.36 (chance-corrected competence).

noul: open 74.92; sealed 74.00 (chance-corrected competence).

score: open 81.49; sealed 74.99 (chance-corrected competence).

0.0095% CI 0.00–0.0077.140.0085.5749.3075.120.04900estimate
101 ≈openJev VerdictHemant (heman10x) · jev-rebuild
Evaluation setup

Source entry: openJev Verdict (heman10x, ModernBERT-base 151M). Published attribution: Hemant (heman10x).

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.402 s; p95 1.018 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 26.78; sealed 16.64 (chance-corrected competence).

noul: open -15.15; sealed -79.16 (chance-corrected competence).

score: open 26.09; sealed 20.26 (chance-corrected competence).

0.0095% CI 0.00–0.010.0052.2283.8886.62-14.090.00279estimate
102 ≈openJev Verdict 1.4Hemant (heman10x) · jev-rebuild
Evaluation setup

Source entry: openJev Verdict 1.4. Published attribution: Hemant (heman10x).

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.771 s; p95 1.120 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 24.29; sealed 17.86 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 31.47; sealed 27.30 (chance-corrected competence).

0.0095% CI 0.00–0.000.0080.3380.6486.62-18.280.00279estimate
103 ≈Qwen3.8-27B (Chutes TEE)Alibaba · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Qwen3.8 27B (Chutes TEE). Published attribution: Qwen / Chutes.

the operator's hosted API

Adjusted latency: p50 6.489 s; p95 31.856 s. none (hosted API, measured as is)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 0.00; C 0.00.

choice: open 97.18; sealed 97.66 (chance-corrected competence).

noul: open 86.68; sealed 98.80 (chance-corrected competence).

score: open 94.83; sealed 98.22 (chance-corrected competence).

0.0095% CI 0.00–0.0095.5698.0956.850.0098.232.17839estimate
104 ≈verdict-smallManavarya09 (Manav) · jev-rebuild
Evaluation setup

Source entry: verdict-small (Manavarya09, multilingual-e5-small 118M). Published attribution: Manavarya09 (Manav).

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.218 s; p95 2.854 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 26.45; sealed 12.71 (chance-corrected competence).

noul: open -55.15; sealed -43.89 (chance-corrected competence).

score: open 25.85; sealed 22.34 (chance-corrected competence).

0.0095% CI 0.00–0.010.0059.4582.06100.00-2.950.00087estimate
105Von 395Mwfzyx (Victor Hugo) · jev-rebuild
Evaluation setup

Source entry: Von (wfzyx, Option-Marker 395M). Published attribution: wfzyx (Victor Hugo).

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.922 s; p95 2.906 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 34.42; sealed 17.86 (chance-corrected competence).

noul: open -78.58; sealed -98.00 (chance-corrected competence).

score: open 38.49; sealed 30.48 (chance-corrected competence).

0.0095% CI 0.00–0.000.0083.4875.7282.71-16.550.00377estimate
106Laya typed-decisionsConvai Innovations · system-one-open
Evaluation setup

Source entry: Laya typed-decisions. Published attribution: Convai Innovations.

Model author or lab

evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1)

Adjusted latency: p50 3.317 s; p95 13.840 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor.

Secondary composites: B 0.00; C 0.00.

choice: open 31.18; sealed 17.84 (chance-corrected competence).

noul: open -91.90; sealed -99.20 (chance-corrected competence).

score: open 35.34; sealed 30.59 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

0.0095% CI 0.00–0.000.0083.3463.3882.53-16.920.00382estimate
Unrankedclassifier.dev (fast)classifier.dev · jev-serviceAPI: sealed item text sent to the operatorruns on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2)
Evaluation setup

Source entry: classifier.dev (fast tier). Published attribution: mrmps (@michael_chomsky).

Model author or lab

the operator's hosted API

Adjusted latency: p50 0.518 s; p95 1.221 s. none (hosted API, measured as is)

ESTIMATE (proxy tokens, I-2): operator list price: higher of 19 Sep plan cost and 26 Sep usage tariff USD 0.042/M input (rule 1.2); no exact base-model floor applies

Secondary composites: B 74.88; C 74.66.

choice: open 84.17; sealed 85.11 (chance-corrected competence).

noul: open 56.78; sealed 64.44 (chance-corrected competence).

score: open 82.50; sealed 81.87 (chance-corrected competence).

74.6695% CI 73.50–75.2075.8189.1981.9958.8977.140.02345estimate
UnrankedSimpleJev Qwen3.6-35B-A3BFeatherless AI · jev-rebuildAPI: sealed item text sent to the operatorPartial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked.
Evaluation setup

Source entry: SimpleJev Qwen3.6-35B-A3B. Published attribution: Featherless AI.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 1.659 s; p95 1.827 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 0.00; C 0.00.

choice: open 16.22; sealed 8.94 (chance-corrected competence).

noul: open -36.29; sealed -42.69 (chance-corrected competence).

score: open -32.01; sealed -35.15 (chance-corrected competence).

0.0095% CI 0.00–0.000.0078.5075.1835.15-22.970.14512estimate
UnrankedDecision-4B (Eval Engine / Chromia)Eval Engine / Chromia · system-one-openPARTIAL / UNRANKED: 1,550 of 1,624 valid answers; 74 context overflows in our evaluator at 2,048 tokens. No official score or rank.
Evaluation setup

Source entry: Decision-4B (Eval Engine / Chromia). Published attribution: Eval Engine / Chromia.

Model author or lab

evaluator-owned H100 GPU

Adjusted latency: p50 0.233 s; p95 0.379 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B.

Secondary composites: B Not measured; C Not measured.

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

Not measuredNot measuredNot measuredNot measuredNot measuredNot measured0.01372estimate
UnrankedJobe Qwen3.5-4BMantisShrimpdev · Not reportedNot measured in JevBench v1.5; no current score or rank.
Evaluation setup

Source entry: Jobe Qwen3.5-4B (frozen). Published attribution: MantisShrimpdev.

Endpoint setup not reported.

Adjusted latency: p50 Not measured s; p95 Not measured s.

Cost basis not reported.

Not measuredNot measuredNot measuredNot measuredNot measuredNot measuredNot measuredUnreported basis
UnrankedMica v0.1 4Bsky7350 · Not reportedThe frozen refusal policy stopped the run after 1,088 of 1,624 rows: 27 documented refusals were mapped to HTTP 422, then three consecutive passthrough HTTP 400 refusals triggered exit 6. The remaining 536 rows have no scores, so this system is not eligible for an official rank.
Evaluation setup

Source entry: mica-v01-4b. Published attribution: unknown.

Model author or lab

Endpoint setup not reported.

Adjusted latency: p50 Not measured s; p95 Not measured s.

Cost basis not reported.

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

Not measuredNot measuredNot measuredNot measuredNot measuredNot measuredNot measuredUnreported basis
UnrankedOpenJev DiffusionGemma 26B-A4B NVFP4razorback16 / Codiv · Not reportedNot measured in JevBench v1.5; no current score or rank.
Evaluation setup

Source entry: OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16). Published attribution: razorback16 / Codiv.

Endpoint setup not reported.

Adjusted latency: p50 Not measured s; p95 Not measured s.

Cost basis not reported.

Not measuredNot measuredNot measuredNot measuredNot measuredNot measuredNot measuredUnreported basis

Read the JevBench method and result notes, or inspect the pinned official results.

JevBench by Florian Standhartinger and contributors. MIT license and copyright notice.

Benchmarks with a different scope

DecisionBench by Hanno Labs. A separate benchmark for typed decisions. Its applied suite and reasoning track stay separate; results are linked at the source rather than imported into this table.

Decision task benchmarks

The seven original benchmark pages include 40 source tables, with full numeric metrics and profiles for identified evaluated configurations. Each page provides a JSON download. Anonymous submissions and conflicting source values are labeled.

These datasets cover different decisions, from sentiment labels to evidence checks. Compare scores only when the sample, label mapping, language, and evaluation setup match. A subset or binary conversion keeps its own result.

WinoGrande

Resolve an ambiguous reference in a sentence.

RAGTruth

Detect unsupported claims against retrieved evidence.

JudgeBench

Choose the objectively better response from a pair.

BBH

Solve challenging reasoning tasks from BIG-Bench.

JevBench

Evaluate decision systems; public-subset accuracy is separate from the full composite.

TabFact

Check whether a table supports or refutes a statement.

ContractNLI

Classify a contract hypothesis and identify supporting evidence.

Circa

Interpret an indirect answer to a yes/no question.

Belebele

Answer passage-comprehension questions across language variants.

TruthfulQA

Test resistance to common misconceptions; a binary conversion is a separate setup.

Choose a system for your own decisions

Match the decision to the test

Check the rubric, input length, answer options, and model configuration. A reranker adapted to pick an option and a native decision model can answer the same question through different serving paths.

Check confidence against errors

Calibration asks whether reported probabilities match observed outcomes. Test an escalation threshold on your own labeled cases before letting a high-confidence answer approve an action or bypass a larger model.

Price the whole decision

Compare cost per completed decision alongside latency and accuracy. JevBench estimates some hosting costs from base-model reference prices and adjusts self-hosted latency. Those assumptions can change which system fits your budget.

Questions

What does Perplexity Decider’s 85.71% measure?

The 85.71% is Perplexity’s reported accuracy across a fixed 7,210-row panel of 11 tasks from September 2026. Task samples have different sizes, and Decider was measured through the Perplexity API. The percentage measures correctness on those samples; it does not establish probability calibration or replace JevBench’s full composite.

What does Perplexity Decider cost?

Perplexity lists pplx-decider-v1-27b at $0.04 per million input tokens, with free output tokens and no per-request fee. State, images, and all questions count toward input usage. Use the returned input-token count to price a completed decision; a short label alone does not determine the request cost.

How are decision models different from reasoning models?

Decision models answer within a defined set of choices or score levels, often with probabilities your code can use. Reasoning models are evaluated on solving broader problems. A general-purpose model can act as a decision system, but its result depends on the prompt, adapter, and serving setup.

Which benchmarks evaluate AI decision models?

JevBench evaluates typed decision systems across intelligence, calibration, speed, and cost. DecisionBench from Hanno Labs publishes task-level accuracy and probability-quality measures, with applied decisions and reasoning in separate tracks. This page mirrors the pinned JevBench release and links to DecisionBench directly; the two scores are not interchangeable.

Does the Decision Models category change overall rankings?

The category groups decision-system evidence for discovery. JevBench retains its official ranking and remains outside overall and category score calculations because its composite includes serving latency, pricing assumptions, and adapters. Model pages keep the evaluated configuration separate from its base model, so scores are not inherited between them.

Related benchmarks