Skip to main content
BenchLM

JevBench

We show this table for reference; we do not rank on it.

Typed-decision evaluation across 1,624 decisions, combining chance-corrected intelligence, probability calibration, latency, and cost.

JevBench by Florian Standhartinger and contributors. Original benchmark · MIT license and copyright notice. The MIT notice applies to JevBench source code and method documentation. Aggregate results come from Benchmark Heaven's public versioned API. Evaluated model weights, external code, upstream datasets, and other third-party data retain their own terms.

JevBench score (option A) on JevBench — v1.5.4 · retrieved September 30, 2026

We mirror the published jevbench score (option a) view for JevBench. Cygnet (73.70) and Winnow-12B Q8 (73.23) are joint leaders in the published table (statistical tie). We do not use these results to rank models overall.

112 systems, 106 rankedDecision ModelsCurrentDisplay onlyUpdated v1.5.4 · retrieved September 30, 2026

Option A · joint leaders (statistical tie). The ≈ marker identifies a published statistical tie with the next row. Missing pairwise markers establish neither a tie nor a separation; 95% intervals below describe individual composite scores.

Official JevBench 1.5 results (112 systems)
RankSystemScoreIntelligenceCalibrationSpeedCostSealed IntelligenceUSD / 1,000 decisions
1 ≈Cygnetblockbrain · system-one-open
Evaluation setup

Source entry: Cygnet (blockbrain, frozen Gemma-4-12B-it). Published attribution: blockbrain.

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.230 s; p95 0.348 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 73.16; C 73.70.

choice: open 82.92; sealed 80.29 (chance-corrected competence).

noul: open 57.86; sealed 55.67 (chance-corrected competence).

score: open 76.60; sealed 73.20 (chance-corrected competence).

73.7095% CI 72.36–74.4671.0987.0190.9756.4369.720.02834estimate
2Winnow-12B Q8Eldan Ring · jev-rebuild
Evaluation setup

Source entry: Winnow-12B Q8. Published attribution: Eldan Ring.

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.340 s; p95 0.715 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 73.47; C 73.23.

choice: open 77.61; sealed 82.04 (chance-corrected competence).

noul: open 58.61; sealed 66.15 (chance-corrected competence).

score: open 81.21; sealed 80.99 (chance-corrected competence).

73.2395% CI 72.02–73.9974.4384.0786.1456.5676.390.02806estimate
3Jev 1.13.0TypeSafe AI · jevAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Jev 1.13.0 (TypeSafe AI). Published attribution: TypeSafe AI.

Model author or lab

the operator's hosted API

Adjusted latency: p50 0.616 s; p95 0.674 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1); no exact base-model floor applies

Secondary composites: B 72.11; C 72.13.

choice: open 85.73; sealed 87.59 (chance-corrected competence).

noul: open 47.75; sealed 48.62 (chance-corrected competence).

score: open 81.16; sealed 81.14 (chance-corrected competence).

72.1395% CI 71.01–72.6172.0088.0383.8154.7372.450.03230estimate
4JevK5 v0.3 4Ballebee · unclassified
Evaluation setup

Source entry: JevK5 v0.3 (4B). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.182 s; p95 0.243 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 68.11; C 63.24.

choice: open 80.25; sealed 74.25 (chance-corrected competence).

noul: open 42.33; sealed 14.98 (chance-corrected competence).

score: open 62.22; sealed 63.59 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

71.9095% CI 69.39–72.9556.2788.3493.5563.0750.940.01702estimate
5Plumb-4Bcrh225 · unclassified
Evaluation setup

Source entry: Plumb-4B (crh225, JevK5 v0.2 + LoRA). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.185 s; p95 0.245 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 67.75; C 62.00.

choice: open 80.43; sealed 73.30 (chance-corrected competence).

noul: open 42.95; sealed 16.58 (chance-corrected competence).

score: open 59.26; sealed 62.57 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

71.5695% CI 69.20–72.7355.8587.4493.4563.0750.810.01702estimate
6 ≈Jev-Omniakhilaaa3 · jev-rebuild
Evaluation setup

Source entry: Jev-Omni (akhilaaa3, Gemma-4-12B merged). Published attribution: akhilaaa3.

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.375 s; p95 0.908 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 71.30; C 71.50.

choice: open 83.37; sealed 78.70 (chance-corrected competence).

noul: open 55.28; sealed 56.73 (chance-corrected competence).

score: open 72.79; sealed 76.07 (chance-corrected competence).

71.5095% CI 70.21–72.4070.4982.6084.6856.0570.500.02917estimate
7decider-4b v2Mapika · system-one-open
Evaluation setup

Source entry: decider-4b v2 (Mapika). Published attribution: Mapika.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.203 s; p95 0.405 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 67.53; C 61.59.

choice: open 80.43; sealed 77.31 (chance-corrected competence).

noul: open 42.92; sealed 18.73 (chance-corrected competence).

score: open 61.88; sealed 53.36 (chance-corrected competence).

71.2895% CI 69.11–72.3555.7785.5990.8664.5449.800.01520estimate
8Decision 4B v1.2FlyMy.AI · unclassified
Evaluation setup

Source entry: Decision 4B v1.2 (FlyMyJev, Qwen3.5-4B + LoRA). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.183 s; p95 0.244 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 66.57; C 56.66.

choice: open 80.87; sealed 73.76 (chance-corrected competence).

noul: open 20.13; sealed 16.44 (chance-corrected competence).

score: open 67.79; sealed 63.00 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

70.8395% CI 68.52–71.9753.6688.5693.5063.0751.060.01702estimate
9Imajev-4B (RTX 5090)mohit67890 · unclassified
Evaluation setup

Source entry: Imajev-4B (RTX 5090). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (RTX 5090), offline read-only container

Adjusted latency: p50 0.234 s; p95 0.329 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE (I-2): measured input tokens; zero generated output tokens for signed logits readout. M2 floor uses the 25 Sep 2026 DeepInfra Qwen3.5-4B snapshot rates (USD 0.03/M input, USD 0.15/M output); the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen3.5-9B.

Secondary composites: B 66.20; C 55.91.

choice: open 80.12; sealed 82.33 (chance-corrected competence).

noul: open 21.82; sealed 7.56 (chance-corrected competence).

score: open 66.80; sealed 62.21 (chance-corrected competence).

v1.5 roster addendum A2. No pairwise comparison is inferred for this addition.

70.3995% CI 67.80–71.6153.4788.1291.1363.2650.700.01677estimate
10Decision 4B v1.1FlyMy.AI · unclassified
Evaluation setup

Source entry: Decision 4B v1.1 (FlyMyJev, Qwen3.5-4B + LoRA). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.183 s; p95 0.243 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 66.09; C 55.14.

choice: open 80.12; sealed 74.77 (chance-corrected competence).

noul: open 18.02; sealed 15.75 (chance-corrected competence).

score: open 65.82; sealed 64.17 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

70.3995% CI 66.86–71.5753.1187.3193.5263.0751.560.01702estimate
11Manchego v2.1oraculumai · system-one-open
Evaluation setup

Source entry: Manchego v2.1. Published attribution: oraculumai.

Model author or lab

evaluator-owned offline RTX6000 GPU

Adjusted latency: p50 0.319 s; p95 0.430 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: frozen 25 Sep DeepInfra Qwen/Qwen3.5-4B market reference, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B). $0.03/M input, zero generated output; 836,500 measured input tokens across 1,624 decisions. Frozen v1.5 base-model floor applied; no bookable Manchego tariff claimed. Re-score if the basis changes.

Secondary composites: B 64.38; C 50.11.

choice: open 66.94; sealed 65.04 (chance-corrected competence).

noul: open 34.26; sealed 16.84 (chance-corrected competence).

score: open 62.03; sealed 62.10 (chance-corrected competence).

v1.5 roster addendum A5. No pairwise comparison is inferred for this addition.

68.8195% CI 59.64–70.2851.2084.9388.6364.3347.990.01545estimate
12 ≈SemIf Qwen3.5-4BTheodore Lee (TheoLeeCJ) · jev-rebuild
Evaluation setup

Source entry: SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ). Published attribution: Theodore Lee (TheoLeeCJ).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.229 s; p95 0.356 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 64.31; C 50.22.

choice: open 73.80; sealed 70.96 (chance-corrected competence).

noul: open 30.30; sealed 13.93 (chance-corrected competence).

score: open 57.89; sealed 61.00 (chance-corrected competence).

68.6695% CI 60.06–70.1651.3183.9790.8963.0748.630.01702estimate
13 ≈spark-s1-4b-v6Abhishek Rai (abhishek085) · jev-rebuild
Evaluation setup

Source entry: spark-s1-4b-v6 (Open Spark Jev, abhishek085). Published attribution: Abhishek Rai (abhishek085).

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.406 s; p95 0.690 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 66.86; C 68.16.

choice: open 72.55; sealed 73.11 (chance-corrected competence).

noul: open 50.44; sealed 46.47 (chance-corrected competence).

score: open 64.39; sealed 65.74 (chance-corrected competence).

68.1695% CI 66.27–69.7062.1269.7185.5360.4261.770.02086estimate
14 ≈metask-jev-4bWayfind (metask-ai) · jev-rebuild
Evaluation setup

Source entry: metask-jev-4b. Published attribution: Wayfind (metask-ai).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.285 s; p95 0.394 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 64.12; C 53.61.

choice: open 70.11; sealed 72.41 (chance-corrected competence).

noul: open 43.23; sealed 25.20 (chance-corrected competence).

score: open 56.88; sealed 53.04 (chance-corrected competence).

67.4895% CI 65.39–68.6953.4882.7489.4957.7350.220.02564estimate
15 ≈HopperHopitAI · jev-rebuild
Evaluation setup

Source entry: Hopper. Published attribution: HopitAI.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.395 s; p95 0.479 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 62.93; C 46.85.

choice: open 70.12; sealed 75.48 (chance-corrected competence).

noul: open 10.23; sealed 12.76 (chance-corrected competence).

score: open 65.76; sealed 64.82 (chance-corrected competence).

67.4795% CI 56.34–69.1549.8687.8787.2362.2651.020.01812estimate
16Malkuth-4Bnewfull5 (dhtocks) · jev-rebuild
Evaluation setup

Source entry: Malkuth-4B (newfull5, Kev post-train). Published attribution: newfull5 (dhtocks).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.442 s; p95 0.548 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 63.88; C 54.99.

choice: open 71.17; sealed 69.55 (chance-corrected competence).

noul: open 39.48; sealed 22.84 (chance-corrected competence).

score: open 64.44; sealed 59.23 (chance-corrected competence).

66.7795% CI 64.94–67.8954.4583.2186.1655.8250.540.02968estimate
17Surogate Rune 26B-A4B v3 (RTX PRO 6000)Surogate · unclassified
Evaluation setup

Source entry: Surogate Rune 26B-A4B v3 (RTX PRO 6000). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (RTX PRO 6000), offline read-only container

Adjusted latency: p50 0.351 s; p95 0.721 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.

Secondary composites: B 66.55; C 66.47.

choice: open 81.99; sealed 82.17 (chance-corrected competence).

noul: open 45.59; sealed 49.53 (chance-corrected competence).

score: open 79.28; sealed 79.88 (chance-corrected competence).

v1.5 roster addendum A2. No pairwise comparison is inferred for this addition.

66.4795% CI 65.43–67.0169.7488.3085.9748.9770.530.05024estimate
18 ≈reflex 4Bkshetrajna12 · jev-rebuild
Evaluation setup

Source entry: reflex 4B (kshetrajna12). Published attribution: kshetrajna12.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 2.865 s; p95 4.627 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 61.93; C 48.05.

choice: open 73.66; sealed 69.31 (chance-corrected competence).

noul: open 33.33; sealed 19.45 (chance-corrected competence).

score: open 58.62; sealed 54.56 (chance-corrected competence).

65.2495% CI 58.56–66.3551.4986.8468.7763.1447.780.01693estimate
19 ≈jev-local Qwen3.5-9Bus (GitHub) · jev-rebuild
Evaluation setup

Source entry: jev-local (Qwen3.5-9B). Published attribution: us (GitHub).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.944 s; p95 5.339 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 63.22; C 57.39.

choice: open 66.88; sealed 65.39 (chance-corrected competence).

noul: open 46.94; sealed 37.02 (chance-corrected competence).

score: open 60.71; sealed 60.71 (chance-corrected competence).

65.2495% CI 63.46–66.5056.2877.8272.9858.8654.370.02352estimate
20 ≈djevMaisa · jev-rebuild
Evaluation setup

Source entry: djev (Maisa, diffusion-gemma). Published attribution: Maisa (David Villalón).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.251 s; p95 0.317 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 64.76; C 64.16.

choice: open 76.62; sealed 72.78 (chance-corrected competence).

noul: open 66.16; sealed 70.25 (chance-corrected competence).

score: open 75.33; sealed 72.80 (chance-corrected competence).

64.1695% CI 63.07–65.0072.3380.4191.0048.2271.950.05321estimate
21 ≈Raw Qwen3 4B Instruct 2507 direct logitsAlibaba · raw-logit-control
Evaluation setup

Source entry: Raw Qwen3 4B Instruct 2507 direct logits. Published attribution: Alibaba Qwen / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.262 s; p95 0.456 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 60.33; C 50.45.

choice: open 65.50; sealed 63.46 (chance-corrected competence).

noul: open 49.81; sealed 38.04 (chance-corrected competence).

score: open 56.98; sealed 50.60 (chance-corrected competence).

62.1495% CI 59.35–64.2254.0652.9589.2363.3650.700.01665estimate
22 ≈jqv Qwen3-32BOctalab · jev-rebuild
Evaluation setup

Source entry: jqv (Qwen3-32B zero-shot). Published attribution: hjmurmur (Octalab).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.307 s; p95 1.503 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 57.50; C 42.20.

choice: open 73.00; sealed 73.09 (chance-corrected competence).

noul: open 18.67; sealed 10.18 (chance-corrected competence).

score: open 62.35; sealed 57.23 (chance-corrected competence).

60.7795% CI 51.25–64.2649.0986.7783.3651.1746.830.04242estimate
23 ≈JevK5 v0.2.0allebee · jev-rebuild
Evaluation setup

Source entry: JevK5 v0.2.0. Published attribution: allebee.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.216 s; p95 0.379 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 53.55; C 40.36.

choice: open 77.61; sealed 71.76 (chance-corrected competence).

noul: open 12.19; sealed -3.89 (chance-corrected competence).

score: open 60.23; sealed 62.33 (chance-corrected competence).

58.1195% CI 47.56–68.2446.7084.8790.8663.0743.400.01702estimate
24Qwen3.5-9B Jev-like data-mix v2jsaurabh · jev-rebuild
Evaluation setup

Source entry: Qwen3.5-9B Jev-like data-mix v2. Published attribution: jsaurabh.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.581 s; p95 1.087 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 52.49; C 53.02.

choice: open 72.91; sealed 74.77 (chance-corrected competence).

noul: open 42.37; sealed 31.67 (chance-corrected competence).

score: open 72.56; sealed 68.41 (chance-corrected competence).

53.0295% CI 51.88–53.8860.4580.8682.0045.6958.280.06461estimate
25 ≈Standard One 8BStandard Thinking · jev-rebuild
Evaluation setup

Source entry: Standard One 8B (Standard Thinking). Published attribution: Standard Thinking (myeongho12).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.203 s; p95 0.278 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 47.13; C 47.15.

choice: open 73.72; sealed 72.18 (chance-corrected competence).

noul: open 45.04; sealed 42.98 (chance-corrected competence).

score: open 61.11; sealed 62.56 (chance-corrected competence).

47.7995% CI 46.67–48.5459.6083.1692.4843.2859.240.07772estimate
26NInfer Qwen3.8-Flash-Next mixedIgor L. / NInfer contributors · native-logit
Evaluation setup

Source entry: NInfer Qwen3.8-Flash-Next mixed. Published attribution: Igor L. / NInfer contributors.

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.310 s; p95 0.445 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 47.70; C 47.47.

choice: open 84.93; sealed 79.11 (chance-corrected competence).

noul: open 51.40; sealed 43.02 (chance-corrected competence).

score: open 74.44; sealed 70.30 (chance-corrected competence).

47.4795% CI 46.70–47.8867.2088.5488.6142.5364.140.08233estimate
27Instinct Dual 4BZooWork · decision-apiAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Instinct Dual 4B. Published attribution: ZooWork / pierre-srp.

Model author or lab

operator-hosted free-preview API

Adjusted latency: p50 0.506 s; p95 1.260 s. x2 (demo assumption, not measured)

ESTIMATE: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule); reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B.

Secondary composites: B 42.97; C 32.62.

choice: open 75.05; sealed 74.49 (chance-corrected competence).

noul: open 4.41; sealed -12.80 (chance-corrected competence).

score: open 59.41; sealed 58.15 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

46.9795% CI 37.86–56.7743.1288.3081.9560.1939.950.02124estimate
28 ≈swanOneblockbrain · system-one-open
Evaluation setup

Source entry: swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4). Published attribution: blockbrain.

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.568 s; p95 0.591 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 47.32; C 46.56.

choice: open 88.11; sealed 82.65 (chance-corrected competence).

noul: open 45.55; sealed 59.89 (chance-corrected competence).

score: open 77.84; sealed 73.31 (chance-corrected competence).

46.5695% CI 45.81–46.9471.2287.0684.7442.1571.950.08478estimate
29 ≈Raw Qwen3 8B direct logitsAlibaba · raw-logit-control
Evaluation setup

Source entry: Raw Qwen3 8B direct logits. Published attribution: Alibaba Qwen / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.333 s; p95 0.621 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 44.63; C 32.85.

choice: open 64.86; sealed 56.55 (chance-corrected competence).

noul: open 48.07; sealed 38.40 (chance-corrected competence).

score: open 57.40; sealed 41.56 (chance-corrected competence).

45.2295% CI 36.89–46.7451.1449.1986.8445.5445.500.06539estimate
30 ≈decider-2bMapika · jev-rebuild
Evaluation setup

Source entry: decider-2b (Mapika). Published attribution: Mapika.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.177 s; p95 0.205 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 41.11; C 31.32.

choice: open 52.02; sealed 53.48 (chance-corrected competence).

noul: open 35.24; sealed 5.78 (chance-corrected competence).

score: open 59.11; sealed 48.39 (chance-corrected competence).

45.1095% CI 33.82–54.2642.3471.5394.4064.9235.890.01477estimate
31 ≈system-one Qwen3-8BSean Goedecke · jev-rebuild
Evaluation setup

Source entry: system-one (Qwen3-8B, Sean Goedecke). Published attribution: Sean Goedecke.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.227 s; p95 0.360 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 43.43; C 31.23.

choice: open 63.18; sealed 56.68 (chance-corrected competence).

noul: open 49.89; sealed 40.04 (chance-corrected competence).

score: open 47.55; sealed 45.51 (chance-corrected competence).

44.1495% CI 36.79–45.6650.4749.3890.8844.9747.410.06829estimate
32 ≈system-one-openmithalouni · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: system-one-open (Gemma 4 E2B LoRA on an L4). Published attribution: mithalouni.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 1.150 s; p95 1.295 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 38.80; C 29.47.

choice: open 59.92; sealed 52.04 (chance-corrected competence).

noul: open 15.40; sealed 14.98 (chance-corrected competence).

score: open 57.99; sealed 49.52 (chance-corrected competence).

42.4495% CI 33.84–51.9241.6471.6378.2768.3338.850.01137estimate
33 ≈Autoloops Gemma 4 31B ITAutoloops · jevAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Autoloops – Gemma 4 31B IT. Published attribution: Autoloops.

the operator's hosted API

Adjusted latency: p50 0.607 s; p95 0.667 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1)

Secondary composites: B 41.84; C 40.53.

choice: open 86.36; sealed 87.51 (chance-corrected competence).

noul: open 68.73; sealed 67.45 (chance-corrected competence).

score: open 75.55; sealed 74.51 (chance-corrected competence).

40.5395% CI 40.09–40.7676.6885.8483.9239.5976.490.10323tariff
34GPT-6 Luna (low)OpenAI · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: GPT-6 Luna (low reasoning effort). Published attribution: OpenAI.

the operator's hosted API

Adjusted latency: p50 1.580 s; p95 3.031 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1); no exact base-model floor applies

Secondary composites: B 43.10; C 40.48.

choice: open 99.06; sealed 97.87 (chance-corrected competence).

noul: open 85.25; sealed 90.18 (chance-corrected competence).

score: open 99.87; sealed 99.52 (chance-corrected competence).

40.4895% CI 40.27–40.6695.2994.9273.2039.0695.860.10750estimate
35 ≈GPT-6 Luna (medium)OpenAI · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: GPT-6 Luna (default medium reasoning effort). Published attribution: OpenAI.

the operator's hosted API

Adjusted latency: p50 1.555 s; p95 3.088 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1); no exact base-model floor applies

Secondary composites: B 41.35; C 38.75.

choice: open 99.38; sealed 99.51 (chance-corrected competence).

noul: open 90.69; sealed 89.93 (chance-corrected competence).

score: open 98.89; sealed 99.02 (chance-corrected competence).

38.7595% CI 38.57–38.9396.2495.5673.1938.3296.150.11379estimate
36 ≈JevOneJuspay · jev-rebuild
Evaluation setup

Source entry: JevOne. Published attribution: Juspay.

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.267 s; p95 0.339 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 37.40; C 31.13.

choice: open 82.43; sealed 77.14 (chance-corrected competence).

noul: open 9.38; sealed 17.13 (chance-corrected competence).

score: open 69.91; sealed 68.81 (chance-corrected competence).

38.2595% CI 37.38–38.7954.1384.9490.4339.8454.360.10122estimate
37 ≈kev 4BJared Palmer · jev-rebuild
Evaluation setup

Source entry: kev 4B (research preview). Published attribution: Jared Palmer.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.493 s; p95 0.574 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 34.59; C 26.44.

choice: open 53.89; sealed 46.76 (chance-corrected competence).

noul: open 29.06; sealed 4.29 (chance-corrected competence).

score: open 58.24; sealed 48.55 (chance-corrected competence).

38.0795% CI 27.48–46.7739.8667.6485.4865.7833.200.01382estimate
38 ≈kev 8BJared Palmer · jev-rebuild
Evaluation setup

Source entry: kev 8B (research preview). Published attribution: Jared Palmer.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.511 s; p95 0.754 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 33.09; C 23.72.

choice: open 60.67; sealed 58.72 (chance-corrected competence).

noul: open 38.20; sealed 14.00 (chance-corrected competence).

score: open 64.04; sealed 54.14 (chance-corrected competence).

34.1595% CI 26.79–37.4848.3071.0284.1440.4242.280.09682estimate
39 ≈open-alternative-jev Qwen3.5-4BIkerMoel · jev-rebuild
Evaluation setup

Source entry: open-alternative-jev (Qwen3.5-4B, IkerMoel). Published attribution: IkerMoel.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.223 s; p95 0.334 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE (proxy tokens, I-2): exact base-model market reference; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 29.92; C 23.31.

choice: open 63.28; sealed 65.00 (chance-corrected competence).

noul: open 5.29; sealed -16.76 (chance-corrected competence).

score: open 59.22; sealed 48.12 (chance-corrected competence).

33.5695% CI 25.75–41.6637.3676.8791.2963.2832.120.01675estimate
40 ≈Bespoke Nimble 9BBespoke Labs · jev-rebuild
Evaluation setup

Source entry: Bespoke Nimble 9B (Bespoke Labs). Published attribution: Bespoke Labs.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.547 s; p95 0.927 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 32.31; C 31.82.

choice: open 71.68; sealed 72.82 (chance-corrected competence).

noul: open 52.86; sealed 44.55 (chance-corrected competence).

score: open 70.90; sealed 69.68 (chance-corrected competence).

31.8295% CI 31.19–32.2763.7577.1682.9536.7562.350.12833estimate
41 ≈Malkuth-2Bnewfull5 (dhtocks) · jev-rebuild
Evaluation setup

Source entry: Malkuth-2B (newfull5, Kev post-train). Published attribution: newfull5 (dhtocks).

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.231 s; p95 0.286 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 26.38; C 20.76.

choice: open 54.83; sealed 48.49 (chance-corrected competence).

noul: open 26.01; sealed -5.78 (chance-corrected competence).

score: open 51.90; sealed 43.12 (chance-corrected competence).

29.9095% CI 20.96–38.9835.5475.1991.8065.5728.610.01405estimate
42openjev-sglang Qwen3.6-35B-A3Bekzhang · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: openjev-sglang (Qwen3.6-35B-A3B on SGLang). Published attribution: ekzhang.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 1.196 s; p95 1.305 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 29.18; C 27.73.

choice: open 80.55; sealed 74.60 (chance-corrected competence).

noul: open 36.38; sealed 27.20 (chance-corrected competence).

score: open 67.14; sealed 66.02 (chance-corrected competence).

29.0395% CI 28.42–29.4058.6582.8778.0735.6355.940.13981estimate
43 ≈decider-35b-a3bMapika · jev-rebuild
Evaluation setup

Source entry: decider-35b-a3b (Mapika). Published attribution: Mapika.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.241 s; p95 0.330 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 27.71; C 27.50.

choice: open 74.30; sealed 76.47 (chance-corrected competence).

noul: open 51.34; sealed 31.31 (chance-corrected competence).

score: open 66.79; sealed 62.55 (chance-corrected competence).

27.5095% CI 26.93–27.8760.4682.0491.0034.3956.780.15387estimate
44local-jev Qwen3.5-4BAmith Chandrappa (amithgc) · jev-rebuild
Evaluation setup

Source entry: local-jev Qwen3.5-4B. Published attribution: Amith Chandrappa (amithgc).

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.427 s; p95 0.991 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 22.73; C 17.93.

choice: open 73.67; sealed 69.06 (chance-corrected competence).

noul: open -23.24; sealed -38.29 (chance-corrected competence).

score: open 59.68; sealed 61.61 (chance-corrected competence).

25.8295% CI 19.89–32.9633.7582.4583.7459.2730.790.02279estimate
45Nemotron Diffusion 8B (optimized vLLM)pst2154 · system-one-open
Evaluation setup

Source entry: Nemotron Diffusion 8B (pst2154, optimized vLLM). Published attribution: pst2154.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.187 s; p95 0.225 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 22.70; C 17.84.

choice: open 57.05; sealed 58.81 (chance-corrected competence).

noul: open 10.77; sealed -16.91 (chance-corrected competence).

score: open 50.79; sealed 42.55 (chance-corrected competence).

v1.5 roster addendum A3. No pairwise comparison is inferred for this addition.

25.6895% CI 18.60–32.7833.8476.2593.7455.4928.150.03045estimate
46 ≈Open-Jev 9BZefan Cai (@Zefan_Cai) · jev-rebuild
Evaluation setup

Source entry: Open-Jev 9B (Zefan Cai). Published attribution: Zefan Cai (@Zefan_Cai).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 1.276 s; p95 3.432 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 24.99; C 24.36.

choice: open 66.19; sealed 71.45 (chance-corrected competence).

noul: open 48.96; sealed 58.69 (chance-corrected competence).

score: open 70.21; sealed 67.42 (chance-corrected competence).

24.3695% CI 23.93–24.6463.8281.5373.5933.0565.860.17041estimate
47 ≈Decision 2B v59FlyMy.AI · jev-rebuild
Evaluation setup

Source entry: Decision 2B (FlyMy.AI, v59). Published attribution: FlyMy.AI (@denti).

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.300 s; p95 0.307 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 19.30; C 15.64.

choice: open 58.27; sealed 67.70 (chance-corrected competence).

noul: open -16.11; sealed -35.13 (chance-corrected competence).

score: open 61.78; sealed 51.36 (chance-corrected competence).

22.5295% CI 16.68–28.7731.3186.2290.3666.3727.980.01321estimate
48 ≈GPT-5.6 Luna (low)OpenAI · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: GPT-5.6 Luna (low reasoning effort). Published attribution: OpenAI.

the operator's hosted API

Adjusted latency: p50 1.321 s; p95 2.970 s. none (hosted API, measured as is)

operator list price

Secondary composites: B 24.16; C 22.37.

choice: open 95.94; sealed 93.85 (chance-corrected competence).

noul: open 87.11; sealed 94.15 (chance-corrected competence).

score: open 96.42; sealed 98.52 (chance-corrected competence).

22.3795% CI 22.26–22.4694.3394.7174.0630.6795.510.20467tariff
49 ≈typecastlmMikhail Gribov · system-one-open
Evaluation setup

Source entry: typecastlm (Mikhail Gribov, Qwen3.5-4B computed head). Published attribution: Mikhail Gribov.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.225 s; p95 0.278 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 18.85; C 15.16.

choice: open 66.45; sealed 72.01 (chance-corrected competence).

noul: open -52.09; sealed -7.49 (chance-corrected competence).

score: open 56.73; sealed 51.83 (chance-corrected competence).

21.8395% CI 16.23–28.5831.2476.7192.0463.9638.780.01590estimate
50 ≈JEV Qwen3.5-9B Base NVFP4WilfLin · jev-rebuild
Evaluation setup

Source entry: JEV Qwen3.5-9B Base NVFP4. Published attribution: WilfLin.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.190 s; p95 0.234 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE (proxy tokens, I-2): exact base-model market reference

Secondary composites: B 17.76; C 13.94.

choice: open 71.17; sealed 59.87 (chance-corrected competence).

noul: open -34.85; sealed -22.47 (chance-corrected competence).

score: open 62.32; sealed 57.43 (chance-corrected competence).

20.0795% CI 15.19–25.5132.2480.9393.5447.5931.610.05584estimate
51Gemini 3.1 Flash-LiteGoogle · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Gemini 3.1 Flash-Lite. Published attribution: Google.

the operator's hosted API

Adjusted latency: p50 0.863 s; p95 1.144 s. none (hosted API, measured as is)

operator list price

Secondary composites: B 20.78; C 19.58.

choice: open 83.98; sealed 84.09 (chance-corrected competence).

noul: open 73.61; sealed 75.05 (chance-corrected competence).

score: open 72.25; sealed 76.53 (chance-corrected competence).

19.5895% CI 19.30–19.8377.5974.6880.0529.7678.560.21941tariff
52AutoJev-27Bdenis-pplx · unclassified
Evaluation setup

Source entry: AutoJev-27B (denis-pplx, Qwen3.8-27B). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.352 s; p95 0.492 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 20.45; C 19.54.

choice: open 85.24; sealed 82.24 (chance-corrected competence).

noul: open 50.44; sealed 62.00 (chance-corrected competence).

score: open 80.06; sealed 76.70 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

19.5495% CI 19.30–19.6372.7887.7087.6229.3773.650.22619estimate
53AutoJev-27B (RTX PRO 6000)denis-pplx · unclassified
Evaluation setup

Source entry: AutoJev-27B (RTX PRO 6000). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (RTX PRO 6000), offline read-only container

Adjusted latency: p50 0.326 s; p95 0.602 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens.

Secondary composites: B 20.39; C 19.48.

choice: open 84.93; sealed 81.70 (chance-corrected competence).

noul: open 51.14; sealed 62.00 (chance-corrected competence).

score: open 79.98; sealed 76.76 (chance-corrected competence).

v1.5 roster addendum A2. No pairwise comparison is inferred for this addition.

19.4895% CI 19.26–19.6072.7586.6887.0729.3773.490.22619estimate
54NInfer Qwen3.8-27B NVFP4Igor L. / NInfer contributors · native-logit
Evaluation setup

Source entry: NInfer Qwen3.8-27B NVFP4. Published attribution: Igor L. / NInfer contributors.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.248 s; p95 0.416 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 19.35; C 18.74.

choice: open 80.86; sealed 80.00 (chance-corrected competence).

noul: open 47.77; sealed 44.73 (chance-corrected competence).

score: open 72.29; sealed 67.19 (chance-corrected competence).

18.7495% CI 18.45–18.9265.4785.9189.8629.1263.970.23052estimate
55Eikos-27Bcaiovicentino1 · unclassified
Evaluation setup

Source entry: Eikos-27B (caiovicentino1, Qwen3.8-27B). Published attribution: unknown.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.354 s; p95 0.488 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 19.48; C 18.50.

choice: open 87.62; sealed 88.27 (chance-corrected competence).

noul: open 56.47; sealed 64.15 (chance-corrected competence).

score: open 78.35; sealed 75.66 (chance-corrected competence).

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

18.5095% CI 18.28–18.6075.0886.2987.6328.6976.020.23824estimate
56 ≈NInfer Qwen3.8-27B NVFP4 (T=1.5)Igor L. / NInfer contributors · native-logit
Evaluation setup

Source entry: NInfer Qwen3.8-27B NVFP4 (T=1.5). Published attribution: Igor L. / NInfer contributors.

Model author or lab

evaluator-owned Lium GPU pod (RTX5090), offline read-only container

Adjusted latency: p50 0.248 s; p95 0.416 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 18.89; C 18.48.

choice: open 80.86; sealed 80.00 (chance-corrected competence).

noul: open 36.83; sealed 36.36 (chance-corrected competence).

score: open 69.61; sealed 62.80 (chance-corrected competence).

18.4895% CI 18.18–18.6561.0886.4989.8629.1259.720.23052estimate
57Instinct Qwen3.8-27BZooWork · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Instinct (ZooWork, Qwen3.8-27B). Published attribution: rayrain-srp (ZooWork).

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 0.519 s; p95 1.308 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 18.86; C 18.32.

choice: open 84.49; sealed 79.67 (chance-corrected competence).

noul: open 28.50; sealed 39.56 (chance-corrected competence).

score: open 72.75; sealed 71.52 (chance-corrected competence).

18.3295% CI 18.03–18.5162.7584.9681.6829.1663.580.22979estimate
58OpenJev (thinking, BF16)razorback16 · jev-rebuild
Evaluation setup

Source entry: OpenJev (thinking, BF16). Published attribution: razorback16.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 1.561 s; p95 2.632 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 19.27; C 17.94.

choice: open 81.67; sealed 78.51 (chance-corrected competence).

noul: open 80.61; sealed 88.62 (chance-corrected competence).

score: open 85.26; sealed 90.57 (chance-corrected competence).

17.9495% CI 17.76–18.0884.2183.1073.8628.5185.900.24146estimate
59djev (thinking)Maisa · jev-rebuild
Evaluation setup

Source entry: djev (thinking). Published attribution: David Villalon / Maisa.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 1.481 s; p95 3.959 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 18.47; C 17.40.

choice: open 64.77; sealed 48.29 (chance-corrected competence).

noul: open 74.63; sealed 95.75 (chance-corrected competence).

score: open 90.27; sealed 90.39 (chance-corrected competence).

17.4095% CI 17.28–17.5177.3595.6672.3228.1378.140.24866estimate
60LitJev Qwen3.8-27BZhengxu Yu · jev-rebuild
Evaluation setup

Source entry: LitJev (Qwen3.8-27B). Published attribution: Zhengxu Yu.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 3.068 s; p95 4.906 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 16.74; C 15.41.

choice: open 80.44; sealed 79.03 (chance-corrected competence).

noul: open 23.49; sealed 33.71 (chance-corrected competence).

score: open 70.26; sealed 63.02 (chance-corrected competence).

16.3195% CI 16.03–16.4858.3284.4568.2228.3658.590.24438estimate
61Bev / Bonsai 27BReza Sayar · system-one-open
Evaluation setup

Source entry: Bev / Bonsai 27B. Published attribution: Reza Sayar.

Model author or lab

evaluator-owned H100 GPU

Adjusted latency: p50 2.191 s; p95 2.433 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate. Frozen market reference qwen/qwen3.8-27b.

Secondary composites: B 15.99; C 12.34.

choice: open 74.11; sealed 76.27 (chance-corrected competence).

noul: open 4.29; sealed 29.49 (chance-corrected competence).

score: open 67.71; sealed 66.61 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

15.7795% CI 15.13–16.0553.0877.6872.7328.2357.460.24676estimate
62 ≈Raw Phi-4 mini direct logitsMicrosoft · raw-logit-control
Evaluation setup

Source entry: Raw Phi-4 mini direct logits. Published attribution: Microsoft / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.281 s; p95 0.431 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 13.11; C 10.57.

choice: open 54.10; sealed 56.67 (chance-corrected competence).

noul: open 6.03; sealed -43.64 (chance-corrected competence).

score: open 54.69; sealed 46.85 (chance-corrected competence).

15.2295% CI 10.26–20.7427.6271.4189.1753.2219.960.03625estimate
63 ≈OpenSourceJev Qwen3.5-4B Q4_K_Msabeel111 · jev-rebuild
Evaluation setup

Source entry: OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp). Published attribution: sabeel111.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 1.389 s; p95 4.041 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3)

Secondary composites: B 11.16; C 9.20.

choice: open 64.38; sealed 51.49 (chance-corrected competence).

noul: open -36.12; sealed -37.24 (chance-corrected competence).

score: open 55.31; sealed 56.79 (chance-corrected competence).

13.2595% CI 8.98–18.3125.7776.0472.5169.1423.680.01069estimate
64reflex-27bkshetrajna12 · jev-rebuild
Evaluation setup

Source entry: reflex-27b (Qwen3.8-27B). Published attribution: kshetrajna12.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 2.748 s; p95 4.379 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 13.77; C 13.19.

choice: open 84.80; sealed 80.30 (chance-corrected competence).

noul: open 30.92; sealed 35.71 (chance-corrected competence).

score: open 73.55; sealed 71.49 (chance-corrected competence).

13.1995% CI 12.99–13.3062.7985.9369.2025.8062.500.29734estimate
65 ≈Open-Jev 2BZefan Cai (@Zefan_Cai) · jev-rebuild
Evaluation setup

Source entry: Open-Jev 2B (Zefan Cai). Published attribution: Zefan Cai (@Zefan_Cai).

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 1.008 s; p95 2.631 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 8.44; C 6.30.

choice: open 51.52; sealed 45.49 (chance-corrected competence).

noul: open 17.34; sealed -8.15 (chance-corrected competence).

score: open 47.50; sealed 47.69 (chance-corrected competence).

9.0795% CI 6.62–11.5433.5673.6275.7633.0528.340.17041estimate
66 ≈GLiNER2 largeFastino · classifier
Evaluation setup

Source entry: GLiNER2 large (Fastino). Published attribution: Fastino.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 1.879 s; p95 17.719 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 7.19; C 5.83.

choice: open 26.57; sealed 28.42 (chance-corrected competence).

noul: open 16.35; sealed -0.62 (chance-corrected competence).

score: open 35.16; sealed 29.07 (chance-corrected competence).

8.4095% CI 5.21–12.3422.4942.3864.7877.5918.960.00558estimate
67 ≈Qwen3-Reranker-4BAlibaba · reranker
Evaluation setup

Source entry: Qwen3-Reranker-4B. Published attribution: Qwen.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.511 s; p95 2.140 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: hosted exact-model reference (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies

Secondary composites: B 5.86; C 4.90.

choice: open 56.70; sealed 45.60 (chance-corrected competence).

noul: open -10.06; sealed -45.93 (chance-corrected competence).

score: open 44.40; sealed 40.42 (chance-corrected competence).

7.0695% CI 4.26–10.7421.0376.2679.6148.4013.360.05249estimate
68 ≈DeepSeek V4.1 FlashDeepSeek · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: DeepSeek V4.1 Flash (thinking default). Published attribution: DeepSeek.

the operator's hosted API

Adjusted latency: p50 1.776 s; p95 6.465 s. none (hosted API, measured as is)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 7.41; C 6.65.

choice: open 96.20; sealed 98.23 (chance-corrected competence).

noul: open 86.17; sealed 96.55 (chance-corrected competence).

score: open 93.16; sealed 91.82 (chance-corrected competence).

6.6595% CI 6.62–6.6693.6996.9269.4019.0995.530.49760estimate
69 ≈SimpleJev Qwen3.5-0.8BFeatherless AI · jev-rebuild
Evaluation setup

Source entry: SimpleJev (Qwen3.5-0.8B, CPU). Published attribution: sabeel111 / Featherless AI.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 7.371 s; p95 16.793 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 3.26; C 2.77.

choice: open 28.28; sealed 27.77 (chance-corrected competence).

noul: open -12.83; sealed 4.25 (chance-corrected competence).

score: open 24.54; sealed 28.47 (chance-corrected competence).

4.0095% CI 2.06–6.6616.7546.6759.0770.3120.160.00976estimate
70 ≈SimpleJev Qwen3.8-27BFeatherless AI · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: SimpleJev Qwen3.8-27B. Published attribution: Featherless AI.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 1.679 s; p95 1.920 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 3.71; C 3.36.

choice: open 85.43; sealed 80.67 (chance-corrected competence).

noul: open 51.10; sealed 60.95 (chance-corrected competence).

score: open 79.20; sealed 79.18 (chance-corrected competence).

3.3695% CI 3.33–3.3772.7587.2074.9214.9073.600.68680estimate
71 ≈decision-machine-1milliseconds.ai · decision-apiAPI: sealed item text sent to the operator
Evaluation setup

Source entry: decision-machine-1 (milliseconds.ai). Published attribution: milliseconds.ai (Baptiste Laget).

Model author or lab

the operator's hosted API

Adjusted latency: p50 0.180 s; p95 0.293 s. none (hosted API, measured as is)

operator standard launch list price (interpretation I-1); no exact base-model floor applies

Secondary composites: B 2.50; C 2.25.

choice: open 53.27; sealed 44.94 (chance-corrected competence).

noul: open -17.55; sealed -67.96 (chance-corrected competence).

score: open 46.78; sealed 38.41 (chance-corrected competence).

3.2495% CI 1.73–5.3914.8281.1892.7756.305.130.02863estimate
72GLiNER2.5 multiFastino · classifier
Evaluation setup

Source entry: GLiNER2.5 multi (Fastino, 287M). Published attribution: Fastino.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 1.370 s; p95 16.069 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 2.10; C 1.89.

choice: open 19.44; sealed 17.24 (chance-corrected competence).

noul: open 0.16; sealed -2.58 (chance-corrected competence).

score: open 27.74; sealed 21.98 (chance-corrected competence).

2.7295% CI 1.20–5.1014.0057.9666.5786.6212.210.00279estimate
73Bosun v3.1 0.6BClause Logic · system-one-open
Evaluation setup

Source entry: Bosun v3.1 0.6B. Published attribution: Clause Logic.

Model author or lab

evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1, local_files_only=True), credential-free (no API key, no HF token)

Adjusted latency: p50 4.037 s; p95 16.515 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-base frozen raw-qwen3-0.6b estimate.

Secondary composites: B 1.92; C 1.74.

choice: open 33.44; sealed 36.69 (chance-corrected competence).

noul: open -4.15; sealed -57.85 (chance-corrected competence).

score: open 40.17; sealed 37.15 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

2.5095% CI 1.15–4.5613.5865.1661.7677.555.330.00560estimate
74GLiNER2.5 baseFastino · classifier
Evaluation setup

Source entry: GLiNER2 (Fastino, gliner2.5-base). Published attribution: Fastino.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.998 s; p95 9.285 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 1.80; C 1.58.

choice: open 24.80; sealed 20.99 (chance-corrected competence).

noul: open 1.50; sealed -10.00 (chance-corrected competence).

score: open 25.42; sealed 18.14 (chance-corrected competence).

2.2795% CI 0.90–4.5013.4835.5770.3386.629.710.00279estimate
75Deem 0.8B v1LibertAI · system-one-open
Evaluation setup

Source entry: Deem 0.8B v1. Published attribution: LibertAI.

Model author or lab

evaluator-owned H100 GPU

Adjusted latency: p50 0.501 s; p95 0.739 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate. Frozen market reference deepinfra:Qwen/Qwen3.5-0.8B.

Secondary composites: B 1.66; C 1.48.

choice: open 28.48; sealed 19.54 (chance-corrected competence).

noul: open 7.67; sealed -28.73 (chance-corrected competence).

score: open 37.90; sealed 19.95 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

2.1495% CI 0.90–4.3413.0238.8884.3280.293.590.00454estimate
76 ≈JevActeinptein · jev-rebuildAPI: sealed item text sent to the operator
Evaluation setup

Source entry: JevAct (einptein, jev1-2b-v2). Published attribution: einptein.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 0.802 s; p95 2.836 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 1.14; C 1.06.

choice: open 31.06; sealed 35.28 (chance-corrected competence).

noul: open -10.65; sealed -55.35 (chance-corrected competence).

score: open 38.74; sealed 30.47 (chance-corrected competence).

1.5295% CI 0.52–3.1811.2462.5476.4368.563.470.01117estimate
77 ≈CLM-8BContrastive-LM · system-one-open
Evaluation setup

Source entry: CLM-8B (Contrastive-LM, clm-latest). Published attribution: Contrastive-LM (Kwok, Kang, Suresh, Saad-Falcon, Pavone, Ré, Mirhoseini).

Model author or lab

evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container

Adjusted latency: p50 0.183 s; p95 0.256 s. x2 + 0.15 s (assumption, not measured)

price floor: base-model reference price applied

Secondary composites: B 1.16; C 1.05.

choice: open 19.36; sealed 8.29 (chance-corrected competence).

noul: open 6.95; sealed -13.49 (chance-corrected competence).

score: open 24.96; sealed 22.52 (chance-corrected competence).

1.5195% CI 0.48–3.1711.4348.4893.2950.505.770.04466estimate
78 ≈kev 0.6BJared Palmer · jev-rebuild
Evaluation setup

Source entry: kev 0.6B (research preview). Published attribution: Jared Palmer.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.422 s; p95 0.450 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.92; C 0.88.

choice: open 33.25; sealed 21.31 (chance-corrected competence).

noul: open 1.96; sealed -57.60 (chance-corrected competence).

score: open 41.18; sealed 31.81 (chance-corrected competence).

1.2695% CI 0.49–2.5010.3467.6587.2280.09-1.490.00461estimate
79 ≈Raw Qwen3 0.6B direct logitsAlibaba · raw-logit-control
Evaluation setup

Source entry: Raw Qwen3 0.6B direct logits. Published attribution: Alibaba Qwen / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.276 s; p95 0.331 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.86; C 0.75.

choice: open 17.12; sealed 23.80 (chance-corrected competence).

noul: open 13.42; sealed -4.25 (chance-corrected competence).

score: open -0.51; sealed 13.79 (chance-corrected competence).

1.0895% CI 0.27–2.6110.5621.4590.3977.5811.110.00559estimate
80 ≈GLiNER2.5 smallFastino · classifier
Evaluation setup

Source entry: GLiNER2.5 small (Fastino, 74M). Published attribution: Fastino.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.472 s; p95 3.891 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.65; C 0.62.

choice: open 15.71; sealed 11.87 (chance-corrected competence).

noul: open 4.31; sealed -34.11 (chance-corrected competence).

score: open 31.47; sealed 27.19 (chance-corrected competence).

0.8995% CI 0.23–2.209.1955.9077.3686.621.650.00279estimate
81 ≈Raw Qwen3 1.7B direct logitsAlibaba · raw-logit-control
Evaluation setup

Source entry: Raw Qwen3 1.7B direct logits. Published attribution: Alibaba Qwen / neutral reproduction.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.282 s; p95 0.337 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.67; C 0.59.

choice: open 41.04; sealed 23.29 (chance-corrected competence).

noul: open 11.66; sealed -3.45 (chance-corrected competence).

score: open -15.90; sealed 1.40 (chance-corrected competence).

0.8595% CI 0.15–2.289.6721.5790.2268.557.080.01118estimate
82 ≈MirrorBluusun · jev-rebuild
Evaluation setup

Source entry: Mirror. Published attribution: Bluusun.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 4.578 s; p95 9.436 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.15; C 0.15.

choice: open 1.64; sealed -9.37 (chance-corrected competence).

noul: open -0.61; sealed -4.55 (chance-corrected competence).

score: open 19.21; sealed 27.26 (chance-corrected competence).

0.2295% CI 0.01–0.905.6043.2163.6489.274.450.00228estimate
83 ≈ZeroEntropy zerank-2ZeroEntropy · reranker
Evaluation setup

Source entry: ZeroEntropy zerank-2. Published attribution: ZeroEntropy.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.473 s; p95 1.984 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies

Secondary composites: B 0.09; C 0.09.

choice: open 56.57; sealed 46.96 (chance-corrected competence).

noul: open -68.07; sealed -85.64 (chance-corrected competence).

score: open 41.40; sealed 37.36 (chance-corrected competence).

0.1395% CI 0.01–0.494.7681.7180.2848.40-0.440.05249estimate
84 ≈jeffLogan Markewich · jev-rebuild
Evaluation setup

Source entry: jeff (Logan Markewich, GLiFormer 400M). Published attribution: Logan Markewich.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 7.137 s; p95 37.170 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.08; C 0.08.

choice: open 33.01; sealed 21.79 (chance-corrected competence).

noul: open -28.09; sealed -60.76 (chance-corrected competence).

score: open 33.21; sealed 28.29 (chance-corrected competence).

0.1295% CI 0.00–0.514.4380.2755.7681.09-3.560.00427estimate
85smalljev semantic-v9Aditya (isHeSatoshi) · jev-rebuild
Evaluation setup

Source entry: smalljev semantic-v9. Published attribution: Aditya (isHeSatoshi).

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.451 s; p95 0.501 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.05; C 0.05.

choice: open 31.14; sealed 18.27 (chance-corrected competence).

noul: open -40.40; sealed -63.82 (chance-corrected competence).

score: open 42.06; sealed 34.99 (chance-corrected competence).

0.0795% CI 0.00–0.383.6673.0986.4660.77-3.520.02030estimate
86Laya multilingualConvai Innovations · system-one-open
Evaluation setup

Source entry: Laya multilingual. Published attribution: Convai Innovations.

Model author or lab

evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1)

Adjusted latency: p50 1.036 s; p95 4.097 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor.

Secondary composites: B 0.01; C 0.01.

choice: open 20.06; sealed 9.13 (chance-corrected competence).

noul: open -35.98; sealed -31.05 (chance-corrected competence).

score: open 24.29; sealed 27.68 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

0.0295% CI 0.00–0.272.3543.5873.7282.161.920.00393estimate
87 ≈OpenDecision ModernBERT-largeDeepan Wadhwa · classifier
Evaluation setup

Source entry: OpenDecision (ModernBERT-large zero-shot). Published attribution: Deepan Wadhwa.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.297 s; p95 0.740 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.01; C 0.01.

choice: open 34.17; sealed 20.04 (chance-corrected competence).

noul: open -47.14; sealed -70.91 (chance-corrected competence).

score: open 41.38; sealed 35.08 (chance-corrected competence).

0.0195% CI 0.00–0.192.0772.5286.5879.11-5.260.00497estimate
88 ≈BAAI bge-reranker-v2-m3BAAI · reranker
Evaluation setup

Source entry: BAAI bge-reranker-v2-m3. Published attribution: BAAI.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.209 s; p95 0.422 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: base-model market reference (deepinfra:BAAI/bge-m3)

Secondary composites: B 0.00; C 0.00.

choice: open 4.41; sealed 4.20 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 32.30; sealed 27.38 (chance-corrected competence).

0.0095% CI 0.00–0.000.0083.3390.5559.33-22.800.02269estimate
89 ≈Certo v1AltSlate Labs · jev-rebuild
Evaluation setup

Source entry: Certo v1 (AltSlate Labs). Published attribution: AltSlate Labs.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.264 s; p95 0.273 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open -1.85; sealed -2.05 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 29.49; sealed 27.25 (chance-corrected competence).

0.0095% CI 0.00–0.000.0087.5291.4396.91-24.930.00127estimate
90 ≈Decision Fast v53aFlyMy.AI · jev-rebuild
Evaluation setup

Source entry: Decision Fast (FlyMy.AI, v53a). Published attribution: FlyMy.AI (@denti).

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.266 s; p95 0.270 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 38.36; sealed 18.49 (chance-corrected competence).

noul: open -59.07; sealed -86.84 (chance-corrected competence).

score: open 42.64; sealed 33.26 (chance-corrected competence).

0.0095% CI 0.00–0.000.0076.0891.4480.09-11.700.00461estimate
91 ≈Alibaba GTE Reranker ModernBERT-baseAlibaba · reranker
Evaluation setup

Source entry: Alibaba GTE Reranker ModernBERT-base. Published attribution: Alibaba-NLP.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.212 s; p95 0.336 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: size-class proxy (deepinfra:thenlper/gte-base); no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 13.02; sealed 2.35 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 31.09; sealed 28.18 (chance-corrected competence).

0.0095% CI 0.00–0.000.0075.2091.4969.27-23.160.01058estimate
92 ≈kev 0.5BJared Palmer · jev-rebuild
Evaluation setup

Source entry: kev 0.5B. Published attribution: Jared Palmer.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.357 s; p95 0.400 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 25.24; sealed 11.19 (chance-corrected competence).

noul: open -48.02; sealed -94.55 (chance-corrected competence).

score: open 33.12; sealed 16.56 (chance-corrected competence).

0.0095% CI 0.00–0.000.0064.9888.4580.09-22.260.00461estimate
93 ≈LayaConvai Innovations · jev-rebuild
Evaluation setup

Source entry: Laya (Convai Innovations, ModernBERT-large 421M). Published attribution: Convai Innovations.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 1.485 s; p95 2.747 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 37.10; sealed 12.77 (chance-corrected competence).

noul: open -75.74; sealed -81.53 (chance-corrected competence).

score: open 37.21; sealed 29.20 (chance-corrected competence).

0.0095% CI 0.00–0.000.0073.7173.8984.89-13.190.00319estimate
94 ≈lev-350mFranck Verrot (franckverrot) · jev-rebuild
Evaluation setup

Source entry: lev-350m (Franck Verrot, LFM2.5-350M). Published attribution: Franck Verrot (franckverrot).

Model author or lab

evaluator-owned Lium GPU pod (RTX6000), offline read-only container

Adjusted latency: p50 0.194 s; p95 0.223 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 33.06; sealed 15.01 (chance-corrected competence).

noul: open -52.20; sealed -100.00 (chance-corrected competence).

score: open 35.72; sealed 30.94 (chance-corrected competence).

0.0095% CI 0.00–0.000.0077.6893.6380.04-18.010.00463estimate
95 ≈Qwen3.5-0.8B Decision ModelMourad Ghafiri · jev-rebuild
Evaluation setup

Source entry: Qwen3.5-0.8B Decision Model (Mourad Ghafiri). Published attribution: Mourad Ghafiri.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 1.308 s; p95 5.023 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate

Secondary composites: B 0.00; C 0.00.

choice: open 34.93; sealed 40.44 (chance-corrected competence).

noul: open -69.95; sealed -86.04 (chance-corrected competence).

score: open 45.21; sealed 34.46 (chance-corrected competence).

0.0095% CI 0.00–0.020.0073.5771.8279.52-3.710.00482estimate
96 ≈Mixedbread mxbai-rerank-base-v2Mixedbread · reranker
Evaluation setup

Source entry: Mixedbread mxbai-rerank-base-v2. Published attribution: Mixedbread.

Model author or lab

evaluator-owned Lium GPU pod (A6000), offline read-only container

Adjusted latency: p50 0.235 s; p95 0.495 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-0.6B); no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 5.84; sealed 0.62 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 31.19; sealed 27.26 (chance-corrected competence).

0.0095% CI 0.00–0.000.0086.6189.3460.34-24.040.02100estimate
97 ≈Needle 3 (2-bit)Cactus Compute · small-tool-model
Evaluation setup

Source entry: Needle 3 (Cactus, 2-bit, local CPU). Published attribution: Cactus Compute.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 135.513 s; p95 285.427 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open -3.57; sealed -8.14 (chance-corrected competence).

noul: open -14.52; sealed -23.60 (chance-corrected competence).

score: open -15.26; sealed -31.60 (chance-corrected competence).

0.0095% CI 0.00–0.000.000.0034.1361.57-21.110.01910estimate
98 ≈Needle 3 (options as tools)Cactus Compute · small-tool-model
Evaluation setup

Source entry: Needle 3, options as tools (post-hoc adapter mode). Published attribution: Cactus Compute.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 58.160 s; p95 136.690 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 10.87; sealed 2.20 (chance-corrected competence).

noul: open -12.38; sealed -30.11 (chance-corrected competence).

score: open -15.85; sealed -31.79 (chance-corrected competence).

0.0095% CI 0.00–0.000.000.0041.0061.57-19.900.01910estimate
99 ≈open-jev-deberta-v3-largeKotoba Labs · jev-rebuild
Evaluation setup

Source entry: open-jev-deberta-v3-large (local CPU). Published attribution: Kotoba Labs.

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 2.933 s; p95 5.097 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 21.00; sealed 16.66 (chance-corrected competence).

noul: open -79.16; sealed -87.38 (chance-corrected competence).

score: open 30.24; sealed 26.89 (chance-corrected competence).

0.0095% CI 0.00–0.000.0077.0668.2577.59-14.610.00558estimate
100 ≈Open Jev JSON CanvasJoshuaSP · jev-rebuild
Evaluation setup

Source entry: Open Jev JSON Canvas (JoshuaSP). Published attribution: JoshuaSP.

Model author or lab

evaluator-owned Lium GPU pod (H100), offline read-only container

Adjusted latency: p50 0.437 s; p95 0.635 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 81.05; sealed 76.36 (chance-corrected competence).

noul: open 74.92; sealed 74.00 (chance-corrected competence).

score: open 81.49; sealed 74.99 (chance-corrected competence).

0.0095% CI 0.00–0.0077.140.0085.5749.3075.120.04900estimate
101 ≈openJev VerdictHemant (heman10x) · jev-rebuild
Evaluation setup

Source entry: openJev Verdict (heman10x, ModernBERT-base 151M). Published attribution: Hemant (heman10x).

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.402 s; p95 1.018 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 26.78; sealed 16.64 (chance-corrected competence).

noul: open -15.15; sealed -79.16 (chance-corrected competence).

score: open 26.09; sealed 20.26 (chance-corrected competence).

0.0095% CI 0.00–0.010.0052.2283.8886.62-14.090.00279estimate
102 ≈openJev Verdict 1.4Hemant (heman10x) · jev-rebuild
Evaluation setup

Source entry: openJev Verdict 1.4. Published attribution: Hemant (heman10x).

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.771 s; p95 1.120 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 24.29; sealed 17.86 (chance-corrected competence).

noul: open -100.00; sealed -100.00 (chance-corrected competence).

score: open 31.47; sealed 27.30 (chance-corrected competence).

0.0095% CI 0.00–0.000.0080.3380.6486.62-18.280.00279estimate
103 ≈Qwen3.8-27B (Chutes TEE)Alibaba · llm-baselineAPI: sealed item text sent to the operator
Evaluation setup

Source entry: Qwen3.8 27B (Chutes TEE). Published attribution: Qwen / Chutes.

the operator's hosted API

Adjusted latency: p50 6.489 s; p95 31.856 s. none (hosted API, measured as is)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 0.00; C 0.00.

choice: open 97.18; sealed 97.66 (chance-corrected competence).

noul: open 86.68; sealed 98.80 (chance-corrected competence).

score: open 94.83; sealed 98.22 (chance-corrected competence).

0.0095% CI 0.00–0.0095.5698.0956.850.0098.232.17839estimate
104 ≈verdict-smallManavarya09 (Manav) · jev-rebuild
Evaluation setup

Source entry: verdict-small (Manavarya09, multilingual-e5-small 118M). Published attribution: Manavarya09 (Manav).

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.218 s; p95 2.854 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 26.45; sealed 12.71 (chance-corrected competence).

noul: open -55.15; sealed -43.89 (chance-corrected competence).

score: open 25.85; sealed 22.34 (chance-corrected competence).

0.0095% CI 0.00–0.010.0059.4582.06100.00-2.950.00087estimate
105Von 395Mwfzyx (Victor Hugo) · jev-rebuild
Evaluation setup

Source entry: Von (wfzyx, Option-Marker 395M). Published attribution: wfzyx (Victor Hugo).

Model author or lab

evaluator-owned CPU container, offline read-only, pinned cores

Adjusted latency: p50 0.922 s; p95 2.906 s. x2 + 0.15 s (assumption, not measured)

documented hosted-model estimate; no exact base-model floor applies

Secondary composites: B 0.00; C 0.00.

choice: open 34.42; sealed 17.86 (chance-corrected competence).

noul: open -78.58; sealed -98.00 (chance-corrected competence).

score: open 38.49; sealed 30.48 (chance-corrected competence).

0.0095% CI 0.00–0.000.0083.4875.7282.71-16.550.00377estimate
106Laya typed-decisionsConvai Innovations · system-one-open
Evaluation setup

Source entry: Laya typed-decisions. Published attribution: Convai Innovations.

Model author or lab

evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1)

Adjusted latency: p50 3.317 s; p95 13.840 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor.

Secondary composites: B 0.00; C 0.00.

choice: open 31.18; sealed 17.84 (chance-corrected competence).

noul: open -91.90; sealed -99.20 (chance-corrected competence).

score: open 35.34; sealed 30.59 (chance-corrected competence).

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

0.0095% CI 0.00–0.000.0083.3463.3882.53-16.920.00382estimate
Unrankedclassifier.dev (fast)classifier.dev · jev-serviceAPI: sealed item text sent to the operatorruns on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2)
Evaluation setup

Source entry: classifier.dev (fast tier). Published attribution: mrmps (@michael_chomsky).

Model author or lab

the operator's hosted API

Adjusted latency: p50 0.518 s; p95 1.221 s. none (hosted API, measured as is)

ESTIMATE (proxy tokens, I-2): operator list price: higher of 19 Sep plan cost and 26 Sep usage tariff USD 0.042/M input (rule 1.2); no exact base-model floor applies

Secondary composites: B 74.88; C 74.66.

choice: open 84.17; sealed 85.11 (chance-corrected competence).

noul: open 56.78; sealed 64.44 (chance-corrected competence).

score: open 82.50; sealed 81.87 (chance-corrected competence).

74.6695% CI 73.50–75.2075.8189.1981.9958.8977.140.02345estimate
UnrankedSimpleJev Qwen3.6-35B-A3BFeatherless AI · jev-rebuildAPI: sealed item text sent to the operatorPartial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked.
Evaluation setup

Source entry: SimpleJev Qwen3.6-35B-A3B. Published attribution: Featherless AI.

Model author or lab

the author's public demo endpoint

Adjusted latency: p50 1.659 s; p95 1.827 s. x2 demo-endpoint adjustment (assumption, not measured)

ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule)

Secondary composites: B 0.00; C 0.00.

choice: open 16.22; sealed 8.94 (chance-corrected competence).

noul: open -36.29; sealed -42.69 (chance-corrected competence).

score: open -32.01; sealed -35.15 (chance-corrected competence).

0.0095% CI 0.00–0.000.0078.5075.1835.15-22.970.14512estimate
UnrankedDecision-4B (Eval Engine / Chromia)Eval Engine / Chromia · system-one-openPARTIAL / UNRANKED: 1,550 of 1,624 valid answers; 74 context overflows in our evaluator at 2,048 tokens. No official score or rank.
Evaluation setup

Source entry: Decision-4B (Eval Engine / Chromia). Published attribution: Eval Engine / Chromia.

Model author or lab

evaluator-owned H100 GPU

Adjusted latency: p50 0.233 s; p95 0.379 s. x2 + 0.15 s (assumption, not measured)

ESTIMATE: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B.

Secondary composites: B Not measured; C Not measured.

v1.5 roster addendum A4. No pairwise comparison is inferred for this addition.

Not measuredNot measuredNot measuredNot measuredNot measuredNot measured0.01372estimate
UnrankedJobe Qwen3.5-4BMantisShrimpdev · Not reportedNot measured in JevBench v1.5; no current score or rank.
Evaluation setup

Source entry: Jobe Qwen3.5-4B (frozen). Published attribution: MantisShrimpdev.

Endpoint setup not reported.

Adjusted latency: p50 Not measured s; p95 Not measured s.

Cost basis not reported.

Not measuredNot measuredNot measuredNot measuredNot measuredNot measuredNot measuredUnreported basis
UnrankedMica v0.1 4Bsky7350 · Not reportedThe frozen refusal policy stopped the run after 1,088 of 1,624 rows: 27 documented refusals were mapped to HTTP 422, then three consecutive passthrough HTTP 400 refusals triggered exit 6. The remaining 536 rows have no scores, so this system is not eligible for an official rank.
Evaluation setup

Source entry: mica-v01-4b. Published attribution: unknown.

Model author or lab

Endpoint setup not reported.

Adjusted latency: p50 Not measured s; p95 Not measured s.

Cost basis not reported.

v1.5 roster addendum A1. No pairwise comparison is inferred for this addition.

Not measuredNot measuredNot measuredNot measuredNot measuredNot measuredNot measuredUnreported basis
UnrankedOpenJev DiffusionGemma 26B-A4B NVFP4razorback16 / Codiv · Not reportedNot measured in JevBench v1.5; no current score or rank.
Evaluation setup

Source entry: OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16). Published attribution: razorback16 / Codiv.

Endpoint setup not reported.

Adjusted latency: p50 Not measured s; p95 Not measured s.

Cost basis not reported.

Not measuredNot measuredNot measuredNot measuredNot measuredNot measuredNot measuredUnreported basis

How to read this leaderboard

The current official option A combines Intelligence, Calibration, Speed, and Cost with equal weights in a harmonic mean, with gates and a generalization penalty. Choice, Noul, and Score each count one third. Intelligence weights open and sealed decisions equally. A statistical tie means the published paired 95% difference interval contains zero.

Operator receipt: 112 sourced rows are currently displayable on this page; the published table names Cygnet and Winnow-12B Q8 as joint leaders (statistical tie).

Honest limit: Florian selected the equal-axis, equal-type headline after reviewing results, as disclosed in the method amendment. Latency includes assumed adjustments, and many costs use frozen base-model estimates. New roster additions have individual intervals without newly inferred pairwise ties. This composite stays outside overall and category scoring; 1.4 and 1.5 are separate protocols.

How to read the JevBench results

JevBench is Benchmark Heaven's evaluation, created by Florian Standhartinger and contributors. Each row names the evaluated model or system and its developer. Reasoning settings, adapters, and serving details describe the configuration that produced the result.

We mirror the complete v1.5.4 roster: 106 ranked systems, one honorable mention, two partial entries, and three incomplete or unmeasured entries. Official ranks and numbers stay attached to the exact configuration. An absent score stays missing; a measured zero stays zero.

The current headline is option A: an equal-weight harmonic mean of Intelligence, Calibration, Speed, and Cost. Choice, Noul, and Score each receive one third. Intelligence gives equal weight to 904 open and 720 sealed decisions. The original frozen method named option B as its headline; Florian selected option A after reviewing results, as disclosed in the headline amendment.

We retain published 95% composite intervals and adjacent-pair statistical ties. Cygnet and Winnow-12B Q8 are joint leaders in the official table. Addendum systems have individual intervals, but missing pairwise markers establish neither a tie nor a separation. The sealed Intelligence column is chance-corrected competence, not percent accuracy.

Cost is USD per 1,000 decisions, with the source tariff or estimate and frozen pricing assumptions retained. It does not establish a token tariff for the evaluated configuration. This composite stays outside overall and category scores. Version 1.4 used different tasks and scoring, so movement between versions is not a like-for-like comparison.

How JevBench changed

This page follows the latest imported release, currently v1.5.4. Earlier results remain labeled by version on model pages. Changes to the tasks and scoring prevent a direct comparison of scores across the 1.4 and 1.5 protocols.

JevBench releases and protocol changes
ReleaseDecisions per complete runWhat changed
v1.5.4 · current imported release904 open + 720 sealedThe roster has 112 systems and 106 official ranks. Manchego v2.1 joins after a complete run; earlier scores, intervals, and the scoring method remain unchanged.
v1.5 · protocol change904 open + 720 sealedChoice, Noul, and Score receive equal weight, and sealed results carry half of Intelligence. The official headline changed to option A after the method owner reviewed results. Published intervals and paired statistical ties help distinguish small gaps.
v1.4.2.2 · September 27, 2026534 frozen + 308 sealedImajev-4B joined a roster of 95 systems, with 91 ranked. Version 1.4 used a harmonic composite and a 20% sealed share of Intelligence; later roster additions preserved earlier measurements.

Cygnet (73.70) and Winnow-12B Q8 (73.23) are joint leaders in the published table (statistical tie). The table retains individual score intervals and published comparisons between adjacent rows.

112 systems, 106 ranked have been evaluated on JevBench. The benchmark falls in the Decision Models category. JevBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About JevBench

Year

2026

Tasks

904 open decisions plus 720 sealed decisions per complete run

Format

Official option A harmonic composite with published uncertainty

Difficulty

Choice, Noul, and Score decisions with sealed generalization tests

We mirror the complete v1.5.4 roster, including six unranked entries, and preserve every published option A score, rank, confidence interval, and adjacent-pair tie. The table keeps configurations separate and displays missing measurements explicitly. The original frozen method and its later headline amendment remain linked beside the data.

Freshness and provenance

Version

JevBench v1.5.4

Refresh cadence

Pinned release

Staleness state

Current

Question availability

601 of 904 open decisions published; 720 sealed decisions private

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does JevBench 1.5 measure?

JevBench 1.5 evaluates typed decisions using Choice, Noul, and Score requests across 904 open and 720 sealed items. Its current official option A combines Intelligence, Calibration, Speed, and Cost equally. The published table includes composite confidence intervals and statistical ties; its score is an index rather than percent accuracy.

Why are some JevBench 1.5 systems unranked?

The v1.5.4 roster contains 112 systems and 106 official ranks. classifier.dev remains an honorable mention because it runs another entrant. Two entries have partial evaluations, and three are incomplete or unmeasured. We retain their published evidence and reasons while leaving unavailable scores and ranks missing.

Are JevBench 1.4 and 1.5 scores comparable?

The versions use different task sets and scoring. Version 1.5 doubles the sealed share of Intelligence to 50%, adds native evaluation of three decision types, and publishes uncertainty. Model evidence rows retain their versioned benchmark keys, and this page explains the earlier releases. A higher score on 1.5 does not by itself establish that a model improved.

Last updated: v1.5.4 · retrieved September 30, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.