JevBench
We show this table for reference; we do not rank on it.
Typed-decision evaluation across 1,624 decisions, combining chance-corrected intelligence, probability calibration, latency, and cost.
JevBench by Florian Standhartinger and contributors. Original benchmark · MIT license and copyright notice. The MIT notice applies to JevBench source code and method documentation. Aggregate results come from Benchmark Heaven's public versioned API. Evaluated model weights, external code, upstream datasets, and other third-party data retain their own terms.
JevBench score (option A) on JevBench — v1.5.4 · retrieved September 30, 2026
We mirror the published jevbench score (option a) view for JevBench. Cygnet (73.70) and Winnow-12B Q8 (73.23) are joint leaders in the published table (statistical tie). We do not use these results to rank models overall.
Cygnet
blockbrain
system-one-open
Winnow-12B Q8
Eldan Ring
jev-rebuild
Jev 1.13.0
TypeSafe AI
jev · API exposure
112 systems, 106 rankedDecision ModelsCurrentDisplay onlyUpdated v1.5.4 · retrieved September 30, 2026
Option A · joint leaders (statistical tie). The ≈ marker identifies a published statistical tie with the next row. Missing pairwise markers establish neither a tie nor a separation; 95% intervals below describe individual composite scores.
| Rank | System | Score | Intelligence | Calibration | Speed | Cost | Sealed Intelligence | USD / 1,000 decisions |
|---|---|---|---|---|---|---|---|---|
| 1 ≈ | Cygnetblockbrain · system-one-openEvaluation setupSource entry: Cygnet (blockbrain, frozen Gemma-4-12B-it). Published attribution: blockbrain. evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.230 s; p95 0.348 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 73.16; C 73.70. choice: open 82.92; sealed 80.29 (chance-corrected competence). noul: open 57.86; sealed 55.67 (chance-corrected competence). score: open 76.60; sealed 73.20 (chance-corrected competence). | 73.7095% CI 72.36–74.46 | 71.09 | 87.01 | 90.97 | 56.43 | 69.72 | 0.02834estimate |
| 2 | Winnow-12B Q8Eldan Ring · jev-rebuildEvaluation setupSource entry: Winnow-12B Q8. Published attribution: Eldan Ring. evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.340 s; p95 0.715 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 73.47; C 73.23. choice: open 77.61; sealed 82.04 (chance-corrected competence). noul: open 58.61; sealed 66.15 (chance-corrected competence). score: open 81.21; sealed 80.99 (chance-corrected competence). | 73.2395% CI 72.02–73.99 | 74.43 | 84.07 | 86.14 | 56.56 | 76.39 | 0.02806estimate |
| 3 | Jev 1.13.0TypeSafe AI · jevAPI: sealed item text sent to the operatorEvaluation setupSource entry: Jev 1.13.0 (TypeSafe AI). Published attribution: TypeSafe AI. the operator's hosted API Adjusted latency: p50 0.616 s; p95 0.674 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1); no exact base-model floor applies Secondary composites: B 72.11; C 72.13. choice: open 85.73; sealed 87.59 (chance-corrected competence). noul: open 47.75; sealed 48.62 (chance-corrected competence). score: open 81.16; sealed 81.14 (chance-corrected competence). | 72.1395% CI 71.01–72.61 | 72.00 | 88.03 | 83.81 | 54.73 | 72.45 | 0.03230estimate |
| 4 | JevK5 v0.3 4Ballebee · unclassifiedEvaluation setupSource entry: JevK5 v0.3 (4B). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.182 s; p95 0.243 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 68.11; C 63.24. choice: open 80.25; sealed 74.25 (chance-corrected competence). noul: open 42.33; sealed 14.98 (chance-corrected competence). score: open 62.22; sealed 63.59 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 71.9095% CI 69.39–72.95 | 56.27 | 88.34 | 93.55 | 63.07 | 50.94 | 0.01702estimate |
| 5 | Plumb-4Bcrh225 · unclassifiedEvaluation setupSource entry: Plumb-4B (crh225, JevK5 v0.2 + LoRA). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.185 s; p95 0.245 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 67.75; C 62.00. choice: open 80.43; sealed 73.30 (chance-corrected competence). noul: open 42.95; sealed 16.58 (chance-corrected competence). score: open 59.26; sealed 62.57 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 71.5695% CI 69.20–72.73 | 55.85 | 87.44 | 93.45 | 63.07 | 50.81 | 0.01702estimate |
| 6 ≈ | Jev-Omniakhilaaa3 · jev-rebuildEvaluation setupSource entry: Jev-Omni (akhilaaa3, Gemma-4-12B merged). Published attribution: akhilaaa3. evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.375 s; p95 0.908 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 71.30; C 71.50. choice: open 83.37; sealed 78.70 (chance-corrected competence). noul: open 55.28; sealed 56.73 (chance-corrected competence). score: open 72.79; sealed 76.07 (chance-corrected competence). | 71.5095% CI 70.21–72.40 | 70.49 | 82.60 | 84.68 | 56.05 | 70.50 | 0.02917estimate |
| 7 | decider-4b v2Mapika · system-one-openEvaluation setupSource entry: decider-4b v2 (Mapika). Published attribution: Mapika. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.203 s; p95 0.405 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 67.53; C 61.59. choice: open 80.43; sealed 77.31 (chance-corrected competence). noul: open 42.92; sealed 18.73 (chance-corrected competence). score: open 61.88; sealed 53.36 (chance-corrected competence). | 71.2895% CI 69.11–72.35 | 55.77 | 85.59 | 90.86 | 64.54 | 49.80 | 0.01520estimate |
| 8 | Decision 4B v1.2FlyMy.AI · unclassifiedEvaluation setupSource entry: Decision 4B v1.2 (FlyMyJev, Qwen3.5-4B + LoRA). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.183 s; p95 0.244 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 66.57; C 56.66. choice: open 80.87; sealed 73.76 (chance-corrected competence). noul: open 20.13; sealed 16.44 (chance-corrected competence). score: open 67.79; sealed 63.00 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 70.8395% CI 68.52–71.97 | 53.66 | 88.56 | 93.50 | 63.07 | 51.06 | 0.01702estimate |
| 9 | Imajev-4B (RTX 5090)mohit67890 · unclassifiedEvaluation setupSource entry: Imajev-4B (RTX 5090). Published attribution: unknown. evaluator-owned Lium GPU pod (RTX 5090), offline read-only container Adjusted latency: p50 0.234 s; p95 0.329 s. x2 + 0.15 s (assumption, not measured) ESTIMATE (I-2): measured input tokens; zero generated output tokens for signed logits readout. M2 floor uses the 25 Sep 2026 DeepInfra Qwen3.5-4B snapshot rates (USD 0.03/M input, USD 0.15/M output); the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen3.5-9B. Secondary composites: B 66.20; C 55.91. choice: open 80.12; sealed 82.33 (chance-corrected competence). noul: open 21.82; sealed 7.56 (chance-corrected competence). score: open 66.80; sealed 62.21 (chance-corrected competence). v1.5 roster addendum A2. No pairwise comparison is inferred for this addition. | 70.3995% CI 67.80–71.61 | 53.47 | 88.12 | 91.13 | 63.26 | 50.70 | 0.01677estimate |
| 10 | Decision 4B v1.1FlyMy.AI · unclassifiedEvaluation setupSource entry: Decision 4B v1.1 (FlyMyJev, Qwen3.5-4B + LoRA). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.183 s; p95 0.243 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 66.09; C 55.14. choice: open 80.12; sealed 74.77 (chance-corrected competence). noul: open 18.02; sealed 15.75 (chance-corrected competence). score: open 65.82; sealed 64.17 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 70.3995% CI 66.86–71.57 | 53.11 | 87.31 | 93.52 | 63.07 | 51.56 | 0.01702estimate |
| 11 | Manchego v2.1oraculumai · system-one-openEvaluation setupSource entry: Manchego v2.1. Published attribution: oraculumai. evaluator-owned offline RTX6000 GPU Adjusted latency: p50 0.319 s; p95 0.430 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: frozen 25 Sep DeepInfra Qwen/Qwen3.5-4B market reference, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B). $0.03/M input, zero generated output; 836,500 measured input tokens across 1,624 decisions. Frozen v1.5 base-model floor applied; no bookable Manchego tariff claimed. Re-score if the basis changes. Secondary composites: B 64.38; C 50.11. choice: open 66.94; sealed 65.04 (chance-corrected competence). noul: open 34.26; sealed 16.84 (chance-corrected competence). score: open 62.03; sealed 62.10 (chance-corrected competence). v1.5 roster addendum A5. No pairwise comparison is inferred for this addition. | 68.8195% CI 59.64–70.28 | 51.20 | 84.93 | 88.63 | 64.33 | 47.99 | 0.01545estimate |
| 12 ≈ | SemIf Qwen3.5-4BTheodore Lee (TheoLeeCJ) · jev-rebuildEvaluation setupSource entry: SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ). Published attribution: Theodore Lee (TheoLeeCJ). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.229 s; p95 0.356 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 64.31; C 50.22. choice: open 73.80; sealed 70.96 (chance-corrected competence). noul: open 30.30; sealed 13.93 (chance-corrected competence). score: open 57.89; sealed 61.00 (chance-corrected competence). | 68.6695% CI 60.06–70.16 | 51.31 | 83.97 | 90.89 | 63.07 | 48.63 | 0.01702estimate |
| 13 ≈ | spark-s1-4b-v6Abhishek Rai (abhishek085) · jev-rebuildEvaluation setupSource entry: spark-s1-4b-v6 (Open Spark Jev, abhishek085). Published attribution: Abhishek Rai (abhishek085). evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.406 s; p95 0.690 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 66.86; C 68.16. choice: open 72.55; sealed 73.11 (chance-corrected competence). noul: open 50.44; sealed 46.47 (chance-corrected competence). score: open 64.39; sealed 65.74 (chance-corrected competence). | 68.1695% CI 66.27–69.70 | 62.12 | 69.71 | 85.53 | 60.42 | 61.77 | 0.02086estimate |
| 14 ≈ | metask-jev-4bWayfind (metask-ai) · jev-rebuildEvaluation setupSource entry: metask-jev-4b. Published attribution: Wayfind (metask-ai). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.285 s; p95 0.394 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 64.12; C 53.61. choice: open 70.11; sealed 72.41 (chance-corrected competence). noul: open 43.23; sealed 25.20 (chance-corrected competence). score: open 56.88; sealed 53.04 (chance-corrected competence). | 67.4895% CI 65.39–68.69 | 53.48 | 82.74 | 89.49 | 57.73 | 50.22 | 0.02564estimate |
| 15 ≈ | HopperHopitAI · jev-rebuildEvaluation setupSource entry: Hopper. Published attribution: HopitAI. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.395 s; p95 0.479 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 62.93; C 46.85. choice: open 70.12; sealed 75.48 (chance-corrected competence). noul: open 10.23; sealed 12.76 (chance-corrected competence). score: open 65.76; sealed 64.82 (chance-corrected competence). | 67.4795% CI 56.34–69.15 | 49.86 | 87.87 | 87.23 | 62.26 | 51.02 | 0.01812estimate |
| 16 | Malkuth-4Bnewfull5 (dhtocks) · jev-rebuildEvaluation setupSource entry: Malkuth-4B (newfull5, Kev post-train). Published attribution: newfull5 (dhtocks). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.442 s; p95 0.548 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 63.88; C 54.99. choice: open 71.17; sealed 69.55 (chance-corrected competence). noul: open 39.48; sealed 22.84 (chance-corrected competence). score: open 64.44; sealed 59.23 (chance-corrected competence). | 66.7795% CI 64.94–67.89 | 54.45 | 83.21 | 86.16 | 55.82 | 50.54 | 0.02968estimate |
| 17 | Surogate Rune 26B-A4B v3 (RTX PRO 6000)Surogate · unclassifiedEvaluation setupSource entry: Surogate Rune 26B-A4B v3 (RTX PRO 6000). Published attribution: unknown. evaluator-owned Lium GPU pod (RTX PRO 6000), offline read-only container Adjusted latency: p50 0.351 s; p95 0.721 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens. Secondary composites: B 66.55; C 66.47. choice: open 81.99; sealed 82.17 (chance-corrected competence). noul: open 45.59; sealed 49.53 (chance-corrected competence). score: open 79.28; sealed 79.88 (chance-corrected competence). v1.5 roster addendum A2. No pairwise comparison is inferred for this addition. | 66.4795% CI 65.43–67.01 | 69.74 | 88.30 | 85.97 | 48.97 | 70.53 | 0.05024estimate |
| 18 ≈ | reflex 4Bkshetrajna12 · jev-rebuildEvaluation setupSource entry: reflex 4B (kshetrajna12). Published attribution: kshetrajna12. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 2.865 s; p95 4.627 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 61.93; C 48.05. choice: open 73.66; sealed 69.31 (chance-corrected competence). noul: open 33.33; sealed 19.45 (chance-corrected competence). score: open 58.62; sealed 54.56 (chance-corrected competence). | 65.2495% CI 58.56–66.35 | 51.49 | 86.84 | 68.77 | 63.14 | 47.78 | 0.01693estimate |
| 19 ≈ | jev-local Qwen3.5-9Bus (GitHub) · jev-rebuildEvaluation setupSource entry: jev-local (Qwen3.5-9B). Published attribution: us (GitHub). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.944 s; p95 5.339 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 63.22; C 57.39. choice: open 66.88; sealed 65.39 (chance-corrected competence). noul: open 46.94; sealed 37.02 (chance-corrected competence). score: open 60.71; sealed 60.71 (chance-corrected competence). | 65.2495% CI 63.46–66.50 | 56.28 | 77.82 | 72.98 | 58.86 | 54.37 | 0.02352estimate |
| 20 ≈ | djevMaisa · jev-rebuildEvaluation setupSource entry: djev (Maisa, diffusion-gemma). Published attribution: Maisa (David Villalón). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.251 s; p95 0.317 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 64.76; C 64.16. choice: open 76.62; sealed 72.78 (chance-corrected competence). noul: open 66.16; sealed 70.25 (chance-corrected competence). score: open 75.33; sealed 72.80 (chance-corrected competence). | 64.1695% CI 63.07–65.00 | 72.33 | 80.41 | 91.00 | 48.22 | 71.95 | 0.05321estimate |
| 21 ≈ | Raw Qwen3 4B Instruct 2507 direct logitsAlibaba · raw-logit-controlEvaluation setupSource entry: Raw Qwen3 4B Instruct 2507 direct logits. Published attribution: Alibaba Qwen / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.262 s; p95 0.456 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 60.33; C 50.45. choice: open 65.50; sealed 63.46 (chance-corrected competence). noul: open 49.81; sealed 38.04 (chance-corrected competence). score: open 56.98; sealed 50.60 (chance-corrected competence). | 62.1495% CI 59.35–64.22 | 54.06 | 52.95 | 89.23 | 63.36 | 50.70 | 0.01665estimate |
| 22 ≈ | jqv Qwen3-32BOctalab · jev-rebuildEvaluation setupSource entry: jqv (Qwen3-32B zero-shot). Published attribution: hjmurmur (Octalab). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.307 s; p95 1.503 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 57.50; C 42.20. choice: open 73.00; sealed 73.09 (chance-corrected competence). noul: open 18.67; sealed 10.18 (chance-corrected competence). score: open 62.35; sealed 57.23 (chance-corrected competence). | 60.7795% CI 51.25–64.26 | 49.09 | 86.77 | 83.36 | 51.17 | 46.83 | 0.04242estimate |
| 23 ≈ | JevK5 v0.2.0allebee · jev-rebuildEvaluation setupSource entry: JevK5 v0.2.0. Published attribution: allebee. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.216 s; p95 0.379 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 53.55; C 40.36. choice: open 77.61; sealed 71.76 (chance-corrected competence). noul: open 12.19; sealed -3.89 (chance-corrected competence). score: open 60.23; sealed 62.33 (chance-corrected competence). | 58.1195% CI 47.56–68.24 | 46.70 | 84.87 | 90.86 | 63.07 | 43.40 | 0.01702estimate |
| 24 | Qwen3.5-9B Jev-like data-mix v2jsaurabh · jev-rebuildEvaluation setupSource entry: Qwen3.5-9B Jev-like data-mix v2. Published attribution: jsaurabh. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.581 s; p95 1.087 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 52.49; C 53.02. choice: open 72.91; sealed 74.77 (chance-corrected competence). noul: open 42.37; sealed 31.67 (chance-corrected competence). score: open 72.56; sealed 68.41 (chance-corrected competence). | 53.0295% CI 51.88–53.88 | 60.45 | 80.86 | 82.00 | 45.69 | 58.28 | 0.06461estimate |
| 25 ≈ | Standard One 8BStandard Thinking · jev-rebuildEvaluation setupSource entry: Standard One 8B (Standard Thinking). Published attribution: Standard Thinking (myeongho12). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.203 s; p95 0.278 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 47.13; C 47.15. choice: open 73.72; sealed 72.18 (chance-corrected competence). noul: open 45.04; sealed 42.98 (chance-corrected competence). score: open 61.11; sealed 62.56 (chance-corrected competence). | 47.7995% CI 46.67–48.54 | 59.60 | 83.16 | 92.48 | 43.28 | 59.24 | 0.07772estimate |
| 26 | NInfer Qwen3.8-Flash-Next mixedIgor L. / NInfer contributors · native-logitEvaluation setupSource entry: NInfer Qwen3.8-Flash-Next mixed. Published attribution: Igor L. / NInfer contributors. evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.310 s; p95 0.445 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 47.70; C 47.47. choice: open 84.93; sealed 79.11 (chance-corrected competence). noul: open 51.40; sealed 43.02 (chance-corrected competence). score: open 74.44; sealed 70.30 (chance-corrected competence). | 47.4795% CI 46.70–47.88 | 67.20 | 88.54 | 88.61 | 42.53 | 64.14 | 0.08233estimate |
| 27 | Instinct Dual 4BZooWork · decision-apiAPI: sealed item text sent to the operatorEvaluation setupSource entry: Instinct Dual 4B. Published attribution: ZooWork / pierre-srp. operator-hosted free-preview API Adjusted latency: p50 0.506 s; p95 1.260 s. x2 (demo assumption, not measured) ESTIMATE: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule); reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B. Secondary composites: B 42.97; C 32.62. choice: open 75.05; sealed 74.49 (chance-corrected competence). noul: open 4.41; sealed -12.80 (chance-corrected competence). score: open 59.41; sealed 58.15 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 46.9795% CI 37.86–56.77 | 43.12 | 88.30 | 81.95 | 60.19 | 39.95 | 0.02124estimate |
| 28 ≈ | swanOneblockbrain · system-one-openEvaluation setupSource entry: swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4). Published attribution: blockbrain. evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.568 s; p95 0.591 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 47.32; C 46.56. choice: open 88.11; sealed 82.65 (chance-corrected competence). noul: open 45.55; sealed 59.89 (chance-corrected competence). score: open 77.84; sealed 73.31 (chance-corrected competence). | 46.5695% CI 45.81–46.94 | 71.22 | 87.06 | 84.74 | 42.15 | 71.95 | 0.08478estimate |
| 29 ≈ | Raw Qwen3 8B direct logitsAlibaba · raw-logit-controlEvaluation setupSource entry: Raw Qwen3 8B direct logits. Published attribution: Alibaba Qwen / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.333 s; p95 0.621 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 44.63; C 32.85. choice: open 64.86; sealed 56.55 (chance-corrected competence). noul: open 48.07; sealed 38.40 (chance-corrected competence). score: open 57.40; sealed 41.56 (chance-corrected competence). | 45.2295% CI 36.89–46.74 | 51.14 | 49.19 | 86.84 | 45.54 | 45.50 | 0.06539estimate |
| 30 ≈ | decider-2bMapika · jev-rebuildEvaluation setupSource entry: decider-2b (Mapika). Published attribution: Mapika. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.177 s; p95 0.205 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 41.11; C 31.32. choice: open 52.02; sealed 53.48 (chance-corrected competence). noul: open 35.24; sealed 5.78 (chance-corrected competence). score: open 59.11; sealed 48.39 (chance-corrected competence). | 45.1095% CI 33.82–54.26 | 42.34 | 71.53 | 94.40 | 64.92 | 35.89 | 0.01477estimate |
| 31 ≈ | system-one Qwen3-8BSean Goedecke · jev-rebuildEvaluation setupSource entry: system-one (Qwen3-8B, Sean Goedecke). Published attribution: Sean Goedecke. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.227 s; p95 0.360 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 43.43; C 31.23. choice: open 63.18; sealed 56.68 (chance-corrected competence). noul: open 49.89; sealed 40.04 (chance-corrected competence). score: open 47.55; sealed 45.51 (chance-corrected competence). | 44.1495% CI 36.79–45.66 | 50.47 | 49.38 | 90.88 | 44.97 | 47.41 | 0.06829estimate |
| 32 ≈ | system-one-openmithalouni · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: system-one-open (Gemma 4 E2B LoRA on an L4). Published attribution: mithalouni. the author's public demo endpoint Adjusted latency: p50 1.150 s; p95 1.295 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 38.80; C 29.47. choice: open 59.92; sealed 52.04 (chance-corrected competence). noul: open 15.40; sealed 14.98 (chance-corrected competence). score: open 57.99; sealed 49.52 (chance-corrected competence). | 42.4495% CI 33.84–51.92 | 41.64 | 71.63 | 78.27 | 68.33 | 38.85 | 0.01137estimate |
| 33 ≈ | Autoloops Gemma 4 31B ITAutoloops · jevAPI: sealed item text sent to the operatorEvaluation setupSource entry: Autoloops – Gemma 4 31B IT. Published attribution: Autoloops. the operator's hosted API Adjusted latency: p50 0.607 s; p95 0.667 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1) Secondary composites: B 41.84; C 40.53. choice: open 86.36; sealed 87.51 (chance-corrected competence). noul: open 68.73; sealed 67.45 (chance-corrected competence). score: open 75.55; sealed 74.51 (chance-corrected competence). | 40.5395% CI 40.09–40.76 | 76.68 | 85.84 | 83.92 | 39.59 | 76.49 | 0.10323tariff |
| 34 | GPT-6 Luna (low)OpenAI · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: GPT-6 Luna (low reasoning effort). Published attribution: OpenAI. the operator's hosted API Adjusted latency: p50 1.580 s; p95 3.031 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1); no exact base-model floor applies Secondary composites: B 43.10; C 40.48. choice: open 99.06; sealed 97.87 (chance-corrected competence). noul: open 85.25; sealed 90.18 (chance-corrected competence). score: open 99.87; sealed 99.52 (chance-corrected competence). | 40.4895% CI 40.27–40.66 | 95.29 | 94.92 | 73.20 | 39.06 | 95.86 | 0.10750estimate |
| 35 ≈ | GPT-6 Luna (medium)OpenAI · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: GPT-6 Luna (default medium reasoning effort). Published attribution: OpenAI. the operator's hosted API Adjusted latency: p50 1.555 s; p95 3.088 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1); no exact base-model floor applies Secondary composites: B 41.35; C 38.75. choice: open 99.38; sealed 99.51 (chance-corrected competence). noul: open 90.69; sealed 89.93 (chance-corrected competence). score: open 98.89; sealed 99.02 (chance-corrected competence). | 38.7595% CI 38.57–38.93 | 96.24 | 95.56 | 73.19 | 38.32 | 96.15 | 0.11379estimate |
| 36 ≈ | JevOneJuspay · jev-rebuildEvaluation setupSource entry: JevOne. Published attribution: Juspay. evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.267 s; p95 0.339 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 37.40; C 31.13. choice: open 82.43; sealed 77.14 (chance-corrected competence). noul: open 9.38; sealed 17.13 (chance-corrected competence). score: open 69.91; sealed 68.81 (chance-corrected competence). | 38.2595% CI 37.38–38.79 | 54.13 | 84.94 | 90.43 | 39.84 | 54.36 | 0.10122estimate |
| 37 ≈ | kev 4BJared Palmer · jev-rebuildEvaluation setupSource entry: kev 4B (research preview). Published attribution: Jared Palmer. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.493 s; p95 0.574 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 34.59; C 26.44. choice: open 53.89; sealed 46.76 (chance-corrected competence). noul: open 29.06; sealed 4.29 (chance-corrected competence). score: open 58.24; sealed 48.55 (chance-corrected competence). | 38.0795% CI 27.48–46.77 | 39.86 | 67.64 | 85.48 | 65.78 | 33.20 | 0.01382estimate |
| 38 ≈ | kev 8BJared Palmer · jev-rebuildEvaluation setupSource entry: kev 8B (research preview). Published attribution: Jared Palmer. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.511 s; p95 0.754 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 33.09; C 23.72. choice: open 60.67; sealed 58.72 (chance-corrected competence). noul: open 38.20; sealed 14.00 (chance-corrected competence). score: open 64.04; sealed 54.14 (chance-corrected competence). | 34.1595% CI 26.79–37.48 | 48.30 | 71.02 | 84.14 | 40.42 | 42.28 | 0.09682estimate |
| 39 ≈ | open-alternative-jev Qwen3.5-4BIkerMoel · jev-rebuildEvaluation setupSource entry: open-alternative-jev (Qwen3.5-4B, IkerMoel). Published attribution: IkerMoel. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.223 s; p95 0.334 s. x2 + 0.15 s (assumption, not measured) ESTIMATE (proxy tokens, I-2): exact base-model market reference; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 29.92; C 23.31. choice: open 63.28; sealed 65.00 (chance-corrected competence). noul: open 5.29; sealed -16.76 (chance-corrected competence). score: open 59.22; sealed 48.12 (chance-corrected competence). | 33.5695% CI 25.75–41.66 | 37.36 | 76.87 | 91.29 | 63.28 | 32.12 | 0.01675estimate |
| 40 ≈ | Bespoke Nimble 9BBespoke Labs · jev-rebuildEvaluation setupSource entry: Bespoke Nimble 9B (Bespoke Labs). Published attribution: Bespoke Labs. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.547 s; p95 0.927 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 32.31; C 31.82. choice: open 71.68; sealed 72.82 (chance-corrected competence). noul: open 52.86; sealed 44.55 (chance-corrected competence). score: open 70.90; sealed 69.68 (chance-corrected competence). | 31.8295% CI 31.19–32.27 | 63.75 | 77.16 | 82.95 | 36.75 | 62.35 | 0.12833estimate |
| 41 ≈ | Malkuth-2Bnewfull5 (dhtocks) · jev-rebuildEvaluation setupSource entry: Malkuth-2B (newfull5, Kev post-train). Published attribution: newfull5 (dhtocks). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.231 s; p95 0.286 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 26.38; C 20.76. choice: open 54.83; sealed 48.49 (chance-corrected competence). noul: open 26.01; sealed -5.78 (chance-corrected competence). score: open 51.90; sealed 43.12 (chance-corrected competence). | 29.9095% CI 20.96–38.98 | 35.54 | 75.19 | 91.80 | 65.57 | 28.61 | 0.01405estimate |
| 42 | openjev-sglang Qwen3.6-35B-A3Bekzhang · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: openjev-sglang (Qwen3.6-35B-A3B on SGLang). Published attribution: ekzhang. the author's public demo endpoint Adjusted latency: p50 1.196 s; p95 1.305 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 29.18; C 27.73. choice: open 80.55; sealed 74.60 (chance-corrected competence). noul: open 36.38; sealed 27.20 (chance-corrected competence). score: open 67.14; sealed 66.02 (chance-corrected competence). | 29.0395% CI 28.42–29.40 | 58.65 | 82.87 | 78.07 | 35.63 | 55.94 | 0.13981estimate |
| 43 ≈ | decider-35b-a3bMapika · jev-rebuildEvaluation setupSource entry: decider-35b-a3b (Mapika). Published attribution: Mapika. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.241 s; p95 0.330 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 27.71; C 27.50. choice: open 74.30; sealed 76.47 (chance-corrected competence). noul: open 51.34; sealed 31.31 (chance-corrected competence). score: open 66.79; sealed 62.55 (chance-corrected competence). | 27.5095% CI 26.93–27.87 | 60.46 | 82.04 | 91.00 | 34.39 | 56.78 | 0.15387estimate |
| 44 | local-jev Qwen3.5-4BAmith Chandrappa (amithgc) · jev-rebuildEvaluation setupSource entry: local-jev Qwen3.5-4B. Published attribution: Amith Chandrappa (amithgc). evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.427 s; p95 0.991 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 22.73; C 17.93. choice: open 73.67; sealed 69.06 (chance-corrected competence). noul: open -23.24; sealed -38.29 (chance-corrected competence). score: open 59.68; sealed 61.61 (chance-corrected competence). | 25.8295% CI 19.89–32.96 | 33.75 | 82.45 | 83.74 | 59.27 | 30.79 | 0.02279estimate |
| 45 | Nemotron Diffusion 8B (optimized vLLM)pst2154 · system-one-openEvaluation setupSource entry: Nemotron Diffusion 8B (pst2154, optimized vLLM). Published attribution: pst2154. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.187 s; p95 0.225 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 22.70; C 17.84. choice: open 57.05; sealed 58.81 (chance-corrected competence). noul: open 10.77; sealed -16.91 (chance-corrected competence). score: open 50.79; sealed 42.55 (chance-corrected competence). v1.5 roster addendum A3. No pairwise comparison is inferred for this addition. | 25.6895% CI 18.60–32.78 | 33.84 | 76.25 | 93.74 | 55.49 | 28.15 | 0.03045estimate |
| 46 ≈ | Open-Jev 9BZefan Cai (@Zefan_Cai) · jev-rebuildEvaluation setupSource entry: Open-Jev 9B (Zefan Cai). Published attribution: Zefan Cai (@Zefan_Cai). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 1.276 s; p95 3.432 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 24.99; C 24.36. choice: open 66.19; sealed 71.45 (chance-corrected competence). noul: open 48.96; sealed 58.69 (chance-corrected competence). score: open 70.21; sealed 67.42 (chance-corrected competence). | 24.3695% CI 23.93–24.64 | 63.82 | 81.53 | 73.59 | 33.05 | 65.86 | 0.17041estimate |
| 47 ≈ | Decision 2B v59FlyMy.AI · jev-rebuildEvaluation setupSource entry: Decision 2B (FlyMy.AI, v59). Published attribution: FlyMy.AI (@denti). evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.300 s; p95 0.307 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 19.30; C 15.64. choice: open 58.27; sealed 67.70 (chance-corrected competence). noul: open -16.11; sealed -35.13 (chance-corrected competence). score: open 61.78; sealed 51.36 (chance-corrected competence). | 22.5295% CI 16.68–28.77 | 31.31 | 86.22 | 90.36 | 66.37 | 27.98 | 0.01321estimate |
| 48 ≈ | GPT-5.6 Luna (low)OpenAI · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: GPT-5.6 Luna (low reasoning effort). Published attribution: OpenAI. the operator's hosted API Adjusted latency: p50 1.321 s; p95 2.970 s. none (hosted API, measured as is) operator list price Secondary composites: B 24.16; C 22.37. choice: open 95.94; sealed 93.85 (chance-corrected competence). noul: open 87.11; sealed 94.15 (chance-corrected competence). score: open 96.42; sealed 98.52 (chance-corrected competence). | 22.3795% CI 22.26–22.46 | 94.33 | 94.71 | 74.06 | 30.67 | 95.51 | 0.20467tariff |
| 49 ≈ | typecastlmMikhail Gribov · system-one-openEvaluation setupSource entry: typecastlm (Mikhail Gribov, Qwen3.5-4B computed head). Published attribution: Mikhail Gribov. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.225 s; p95 0.278 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 18.85; C 15.16. choice: open 66.45; sealed 72.01 (chance-corrected competence). noul: open -52.09; sealed -7.49 (chance-corrected competence). score: open 56.73; sealed 51.83 (chance-corrected competence). | 21.8395% CI 16.23–28.58 | 31.24 | 76.71 | 92.04 | 63.96 | 38.78 | 0.01590estimate |
| 50 ≈ | JEV Qwen3.5-9B Base NVFP4WilfLin · jev-rebuildEvaluation setupSource entry: JEV Qwen3.5-9B Base NVFP4. Published attribution: WilfLin. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.190 s; p95 0.234 s. x2 + 0.15 s (assumption, not measured) ESTIMATE (proxy tokens, I-2): exact base-model market reference Secondary composites: B 17.76; C 13.94. choice: open 71.17; sealed 59.87 (chance-corrected competence). noul: open -34.85; sealed -22.47 (chance-corrected competence). score: open 62.32; sealed 57.43 (chance-corrected competence). | 20.0795% CI 15.19–25.51 | 32.24 | 80.93 | 93.54 | 47.59 | 31.61 | 0.05584estimate |
| 51 | Gemini 3.1 Flash-LiteGoogle · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: Gemini 3.1 Flash-Lite. Published attribution: Google. the operator's hosted API Adjusted latency: p50 0.863 s; p95 1.144 s. none (hosted API, measured as is) operator list price Secondary composites: B 20.78; C 19.58. choice: open 83.98; sealed 84.09 (chance-corrected competence). noul: open 73.61; sealed 75.05 (chance-corrected competence). score: open 72.25; sealed 76.53 (chance-corrected competence). | 19.5895% CI 19.30–19.83 | 77.59 | 74.68 | 80.05 | 29.76 | 78.56 | 0.21941tariff |
| 52 | AutoJev-27Bdenis-pplx · unclassifiedEvaluation setupSource entry: AutoJev-27B (denis-pplx, Qwen3.8-27B). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.352 s; p95 0.492 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 20.45; C 19.54. choice: open 85.24; sealed 82.24 (chance-corrected competence). noul: open 50.44; sealed 62.00 (chance-corrected competence). score: open 80.06; sealed 76.70 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 19.5495% CI 19.30–19.63 | 72.78 | 87.70 | 87.62 | 29.37 | 73.65 | 0.22619estimate |
| 53 | AutoJev-27B (RTX PRO 6000)denis-pplx · unclassifiedEvaluation setupSource entry: AutoJev-27B (RTX PRO 6000). Published attribution: unknown. evaluator-owned Lium GPU pod (RTX PRO 6000), offline read-only container Adjusted latency: p50 0.326 s; p95 0.602 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens. Secondary composites: B 20.39; C 19.48. choice: open 84.93; sealed 81.70 (chance-corrected competence). noul: open 51.14; sealed 62.00 (chance-corrected competence). score: open 79.98; sealed 76.76 (chance-corrected competence). v1.5 roster addendum A2. No pairwise comparison is inferred for this addition. | 19.4895% CI 19.26–19.60 | 72.75 | 86.68 | 87.07 | 29.37 | 73.49 | 0.22619estimate |
| 54 | NInfer Qwen3.8-27B NVFP4Igor L. / NInfer contributors · native-logitEvaluation setupSource entry: NInfer Qwen3.8-27B NVFP4. Published attribution: Igor L. / NInfer contributors. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.248 s; p95 0.416 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 19.35; C 18.74. choice: open 80.86; sealed 80.00 (chance-corrected competence). noul: open 47.77; sealed 44.73 (chance-corrected competence). score: open 72.29; sealed 67.19 (chance-corrected competence). | 18.7495% CI 18.45–18.92 | 65.47 | 85.91 | 89.86 | 29.12 | 63.97 | 0.23052estimate |
| 55 | Eikos-27Bcaiovicentino1 · unclassifiedEvaluation setupSource entry: Eikos-27B (caiovicentino1, Qwen3.8-27B). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.354 s; p95 0.488 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 19.48; C 18.50. choice: open 87.62; sealed 88.27 (chance-corrected competence). noul: open 56.47; sealed 64.15 (chance-corrected competence). score: open 78.35; sealed 75.66 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 18.5095% CI 18.28–18.60 | 75.08 | 86.29 | 87.63 | 28.69 | 76.02 | 0.23824estimate |
| 56 ≈ | NInfer Qwen3.8-27B NVFP4 (T=1.5)Igor L. / NInfer contributors · native-logitEvaluation setupSource entry: NInfer Qwen3.8-27B NVFP4 (T=1.5). Published attribution: Igor L. / NInfer contributors. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.248 s; p95 0.416 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 18.89; C 18.48. choice: open 80.86; sealed 80.00 (chance-corrected competence). noul: open 36.83; sealed 36.36 (chance-corrected competence). score: open 69.61; sealed 62.80 (chance-corrected competence). | 18.4895% CI 18.18–18.65 | 61.08 | 86.49 | 89.86 | 29.12 | 59.72 | 0.23052estimate |
| 57 | Instinct Qwen3.8-27BZooWork · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: Instinct (ZooWork, Qwen3.8-27B). Published attribution: rayrain-srp (ZooWork). the author's public demo endpoint Adjusted latency: p50 0.519 s; p95 1.308 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 18.86; C 18.32. choice: open 84.49; sealed 79.67 (chance-corrected competence). noul: open 28.50; sealed 39.56 (chance-corrected competence). score: open 72.75; sealed 71.52 (chance-corrected competence). | 18.3295% CI 18.03–18.51 | 62.75 | 84.96 | 81.68 | 29.16 | 63.58 | 0.22979estimate |
| 58 | OpenJev (thinking, BF16)razorback16 · jev-rebuildEvaluation setupSource entry: OpenJev (thinking, BF16). Published attribution: razorback16. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 1.561 s; p95 2.632 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 19.27; C 17.94. choice: open 81.67; sealed 78.51 (chance-corrected competence). noul: open 80.61; sealed 88.62 (chance-corrected competence). score: open 85.26; sealed 90.57 (chance-corrected competence). | 17.9495% CI 17.76–18.08 | 84.21 | 83.10 | 73.86 | 28.51 | 85.90 | 0.24146estimate |
| 59 | djev (thinking)Maisa · jev-rebuildEvaluation setupSource entry: djev (thinking). Published attribution: David Villalon / Maisa. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 1.481 s; p95 3.959 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 18.47; C 17.40. choice: open 64.77; sealed 48.29 (chance-corrected competence). noul: open 74.63; sealed 95.75 (chance-corrected competence). score: open 90.27; sealed 90.39 (chance-corrected competence). | 17.4095% CI 17.28–17.51 | 77.35 | 95.66 | 72.32 | 28.13 | 78.14 | 0.24866estimate |
| 60 | LitJev Qwen3.8-27BZhengxu Yu · jev-rebuildEvaluation setupSource entry: LitJev (Qwen3.8-27B). Published attribution: Zhengxu Yu. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 3.068 s; p95 4.906 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 16.74; C 15.41. choice: open 80.44; sealed 79.03 (chance-corrected competence). noul: open 23.49; sealed 33.71 (chance-corrected competence). score: open 70.26; sealed 63.02 (chance-corrected competence). | 16.3195% CI 16.03–16.48 | 58.32 | 84.45 | 68.22 | 28.36 | 58.59 | 0.24438estimate |
| 61 | Bev / Bonsai 27BReza Sayar · system-one-openEvaluation setupSource entry: Bev / Bonsai 27B. Published attribution: Reza Sayar. evaluator-owned H100 GPU Adjusted latency: p50 2.191 s; p95 2.433 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate. Frozen market reference qwen/qwen3.8-27b. Secondary composites: B 15.99; C 12.34. choice: open 74.11; sealed 76.27 (chance-corrected competence). noul: open 4.29; sealed 29.49 (chance-corrected competence). score: open 67.71; sealed 66.61 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 15.7795% CI 15.13–16.05 | 53.08 | 77.68 | 72.73 | 28.23 | 57.46 | 0.24676estimate |
| 62 ≈ | Raw Phi-4 mini direct logitsMicrosoft · raw-logit-controlEvaluation setupSource entry: Raw Phi-4 mini direct logits. Published attribution: Microsoft / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.281 s; p95 0.431 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 13.11; C 10.57. choice: open 54.10; sealed 56.67 (chance-corrected competence). noul: open 6.03; sealed -43.64 (chance-corrected competence). score: open 54.69; sealed 46.85 (chance-corrected competence). | 15.2295% CI 10.26–20.74 | 27.62 | 71.41 | 89.17 | 53.22 | 19.96 | 0.03625estimate |
| 63 ≈ | OpenSourceJev Qwen3.5-4B Q4_K_Msabeel111 · jev-rebuildEvaluation setupSource entry: OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp). Published attribution: sabeel111. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 1.389 s; p95 4.041 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 11.16; C 9.20. choice: open 64.38; sealed 51.49 (chance-corrected competence). noul: open -36.12; sealed -37.24 (chance-corrected competence). score: open 55.31; sealed 56.79 (chance-corrected competence). | 13.2595% CI 8.98–18.31 | 25.77 | 76.04 | 72.51 | 69.14 | 23.68 | 0.01069estimate |
| 64 | reflex-27bkshetrajna12 · jev-rebuildEvaluation setupSource entry: reflex-27b (Qwen3.8-27B). Published attribution: kshetrajna12. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 2.748 s; p95 4.379 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 13.77; C 13.19. choice: open 84.80; sealed 80.30 (chance-corrected competence). noul: open 30.92; sealed 35.71 (chance-corrected competence). score: open 73.55; sealed 71.49 (chance-corrected competence). | 13.1995% CI 12.99–13.30 | 62.79 | 85.93 | 69.20 | 25.80 | 62.50 | 0.29734estimate |
| 65 ≈ | Open-Jev 2BZefan Cai (@Zefan_Cai) · jev-rebuildEvaluation setupSource entry: Open-Jev 2B (Zefan Cai). Published attribution: Zefan Cai (@Zefan_Cai). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 1.008 s; p95 2.631 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 8.44; C 6.30. choice: open 51.52; sealed 45.49 (chance-corrected competence). noul: open 17.34; sealed -8.15 (chance-corrected competence). score: open 47.50; sealed 47.69 (chance-corrected competence). | 9.0795% CI 6.62–11.54 | 33.56 | 73.62 | 75.76 | 33.05 | 28.34 | 0.17041estimate |
| 66 ≈ | GLiNER2 largeFastino · classifierEvaluation setupSource entry: GLiNER2 large (Fastino). Published attribution: Fastino. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 1.879 s; p95 17.719 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 7.19; C 5.83. choice: open 26.57; sealed 28.42 (chance-corrected competence). noul: open 16.35; sealed -0.62 (chance-corrected competence). score: open 35.16; sealed 29.07 (chance-corrected competence). | 8.4095% CI 5.21–12.34 | 22.49 | 42.38 | 64.78 | 77.59 | 18.96 | 0.00558estimate |
| 67 ≈ | Qwen3-Reranker-4BAlibaba · rerankerEvaluation setupSource entry: Qwen3-Reranker-4B. Published attribution: Qwen. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.511 s; p95 2.140 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: hosted exact-model reference (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies Secondary composites: B 5.86; C 4.90. choice: open 56.70; sealed 45.60 (chance-corrected competence). noul: open -10.06; sealed -45.93 (chance-corrected competence). score: open 44.40; sealed 40.42 (chance-corrected competence). | 7.0695% CI 4.26–10.74 | 21.03 | 76.26 | 79.61 | 48.40 | 13.36 | 0.05249estimate |
| 68 ≈ | DeepSeek V4.1 FlashDeepSeek · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: DeepSeek V4.1 Flash (thinking default). Published attribution: DeepSeek. the operator's hosted API Adjusted latency: p50 1.776 s; p95 6.465 s. none (hosted API, measured as is) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 7.41; C 6.65. choice: open 96.20; sealed 98.23 (chance-corrected competence). noul: open 86.17; sealed 96.55 (chance-corrected competence). score: open 93.16; sealed 91.82 (chance-corrected competence). | 6.6595% CI 6.62–6.66 | 93.69 | 96.92 | 69.40 | 19.09 | 95.53 | 0.49760estimate |
| 69 ≈ | SimpleJev Qwen3.5-0.8BFeatherless AI · jev-rebuildEvaluation setupSource entry: SimpleJev (Qwen3.5-0.8B, CPU). Published attribution: sabeel111 / Featherless AI. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 7.371 s; p95 16.793 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 3.26; C 2.77. choice: open 28.28; sealed 27.77 (chance-corrected competence). noul: open -12.83; sealed 4.25 (chance-corrected competence). score: open 24.54; sealed 28.47 (chance-corrected competence). | 4.0095% CI 2.06–6.66 | 16.75 | 46.67 | 59.07 | 70.31 | 20.16 | 0.00976estimate |
| 70 ≈ | SimpleJev Qwen3.8-27BFeatherless AI · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: SimpleJev Qwen3.8-27B. Published attribution: Featherless AI. the author's public demo endpoint Adjusted latency: p50 1.679 s; p95 1.920 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 3.71; C 3.36. choice: open 85.43; sealed 80.67 (chance-corrected competence). noul: open 51.10; sealed 60.95 (chance-corrected competence). score: open 79.20; sealed 79.18 (chance-corrected competence). | 3.3695% CI 3.33–3.37 | 72.75 | 87.20 | 74.92 | 14.90 | 73.60 | 0.68680estimate |
| 71 ≈ | decision-machine-1milliseconds.ai · decision-apiAPI: sealed item text sent to the operatorEvaluation setupSource entry: decision-machine-1 (milliseconds.ai). Published attribution: milliseconds.ai (Baptiste Laget). the operator's hosted API Adjusted latency: p50 0.180 s; p95 0.293 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1); no exact base-model floor applies Secondary composites: B 2.50; C 2.25. choice: open 53.27; sealed 44.94 (chance-corrected competence). noul: open -17.55; sealed -67.96 (chance-corrected competence). score: open 46.78; sealed 38.41 (chance-corrected competence). | 3.2495% CI 1.73–5.39 | 14.82 | 81.18 | 92.77 | 56.30 | 5.13 | 0.02863estimate |
| 72 | GLiNER2.5 multiFastino · classifierEvaluation setupSource entry: GLiNER2.5 multi (Fastino, 287M). Published attribution: Fastino. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 1.370 s; p95 16.069 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 2.10; C 1.89. choice: open 19.44; sealed 17.24 (chance-corrected competence). noul: open 0.16; sealed -2.58 (chance-corrected competence). score: open 27.74; sealed 21.98 (chance-corrected competence). | 2.7295% CI 1.20–5.10 | 14.00 | 57.96 | 66.57 | 86.62 | 12.21 | 0.00279estimate |
| 73 | Bosun v3.1 0.6BClause Logic · system-one-openEvaluation setupSource entry: Bosun v3.1 0.6B. Published attribution: Clause Logic. evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1, local_files_only=True), credential-free (no API key, no HF token) Adjusted latency: p50 4.037 s; p95 16.515 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-base frozen raw-qwen3-0.6b estimate. Secondary composites: B 1.92; C 1.74. choice: open 33.44; sealed 36.69 (chance-corrected competence). noul: open -4.15; sealed -57.85 (chance-corrected competence). score: open 40.17; sealed 37.15 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 2.5095% CI 1.15–4.56 | 13.58 | 65.16 | 61.76 | 77.55 | 5.33 | 0.00560estimate |
| 74 | GLiNER2.5 baseFastino · classifierEvaluation setupSource entry: GLiNER2 (Fastino, gliner2.5-base). Published attribution: Fastino. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.998 s; p95 9.285 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 1.80; C 1.58. choice: open 24.80; sealed 20.99 (chance-corrected competence). noul: open 1.50; sealed -10.00 (chance-corrected competence). score: open 25.42; sealed 18.14 (chance-corrected competence). | 2.2795% CI 0.90–4.50 | 13.48 | 35.57 | 70.33 | 86.62 | 9.71 | 0.00279estimate |
| 75 | Deem 0.8B v1LibertAI · system-one-openEvaluation setupSource entry: Deem 0.8B v1. Published attribution: LibertAI. evaluator-owned H100 GPU Adjusted latency: p50 0.501 s; p95 0.739 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate. Frozen market reference deepinfra:Qwen/Qwen3.5-0.8B. Secondary composites: B 1.66; C 1.48. choice: open 28.48; sealed 19.54 (chance-corrected competence). noul: open 7.67; sealed -28.73 (chance-corrected competence). score: open 37.90; sealed 19.95 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 2.1495% CI 0.90–4.34 | 13.02 | 38.88 | 84.32 | 80.29 | 3.59 | 0.00454estimate |
| 76 ≈ | JevActeinptein · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: JevAct (einptein, jev1-2b-v2). Published attribution: einptein. the author's public demo endpoint Adjusted latency: p50 0.802 s; p95 2.836 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 1.14; C 1.06. choice: open 31.06; sealed 35.28 (chance-corrected competence). noul: open -10.65; sealed -55.35 (chance-corrected competence). score: open 38.74; sealed 30.47 (chance-corrected competence). | 1.5295% CI 0.52–3.18 | 11.24 | 62.54 | 76.43 | 68.56 | 3.47 | 0.01117estimate |
| 77 ≈ | CLM-8BContrastive-LM · system-one-openEvaluation setupSource entry: CLM-8B (Contrastive-LM, clm-latest). Published attribution: Contrastive-LM (Kwok, Kang, Suresh, Saad-Falcon, Pavone, Ré, Mirhoseini). evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.183 s; p95 0.256 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 1.16; C 1.05. choice: open 19.36; sealed 8.29 (chance-corrected competence). noul: open 6.95; sealed -13.49 (chance-corrected competence). score: open 24.96; sealed 22.52 (chance-corrected competence). | 1.5195% CI 0.48–3.17 | 11.43 | 48.48 | 93.29 | 50.50 | 5.77 | 0.04466estimate |
| 78 ≈ | kev 0.6BJared Palmer · jev-rebuildEvaluation setupSource entry: kev 0.6B (research preview). Published attribution: Jared Palmer. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.422 s; p95 0.450 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.92; C 0.88. choice: open 33.25; sealed 21.31 (chance-corrected competence). noul: open 1.96; sealed -57.60 (chance-corrected competence). score: open 41.18; sealed 31.81 (chance-corrected competence). | 1.2695% CI 0.49–2.50 | 10.34 | 67.65 | 87.22 | 80.09 | -1.49 | 0.00461estimate |
| 79 ≈ | Raw Qwen3 0.6B direct logitsAlibaba · raw-logit-controlEvaluation setupSource entry: Raw Qwen3 0.6B direct logits. Published attribution: Alibaba Qwen / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.276 s; p95 0.331 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.86; C 0.75. choice: open 17.12; sealed 23.80 (chance-corrected competence). noul: open 13.42; sealed -4.25 (chance-corrected competence). score: open -0.51; sealed 13.79 (chance-corrected competence). | 1.0895% CI 0.27–2.61 | 10.56 | 21.45 | 90.39 | 77.58 | 11.11 | 0.00559estimate |
| 80 ≈ | GLiNER2.5 smallFastino · classifierEvaluation setupSource entry: GLiNER2.5 small (Fastino, 74M). Published attribution: Fastino. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.472 s; p95 3.891 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.65; C 0.62. choice: open 15.71; sealed 11.87 (chance-corrected competence). noul: open 4.31; sealed -34.11 (chance-corrected competence). score: open 31.47; sealed 27.19 (chance-corrected competence). | 0.8995% CI 0.23–2.20 | 9.19 | 55.90 | 77.36 | 86.62 | 1.65 | 0.00279estimate |
| 81 ≈ | Raw Qwen3 1.7B direct logitsAlibaba · raw-logit-controlEvaluation setupSource entry: Raw Qwen3 1.7B direct logits. Published attribution: Alibaba Qwen / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.282 s; p95 0.337 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.67; C 0.59. choice: open 41.04; sealed 23.29 (chance-corrected competence). noul: open 11.66; sealed -3.45 (chance-corrected competence). score: open -15.90; sealed 1.40 (chance-corrected competence). | 0.8595% CI 0.15–2.28 | 9.67 | 21.57 | 90.22 | 68.55 | 7.08 | 0.01118estimate |
| 82 ≈ | MirrorBluusun · jev-rebuildEvaluation setupSource entry: Mirror. Published attribution: Bluusun. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 4.578 s; p95 9.436 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.15; C 0.15. choice: open 1.64; sealed -9.37 (chance-corrected competence). noul: open -0.61; sealed -4.55 (chance-corrected competence). score: open 19.21; sealed 27.26 (chance-corrected competence). | 0.2295% CI 0.01–0.90 | 5.60 | 43.21 | 63.64 | 89.27 | 4.45 | 0.00228estimate |
| 83 ≈ | ZeroEntropy zerank-2ZeroEntropy · rerankerEvaluation setupSource entry: ZeroEntropy zerank-2. Published attribution: ZeroEntropy. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.473 s; p95 1.984 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies Secondary composites: B 0.09; C 0.09. choice: open 56.57; sealed 46.96 (chance-corrected competence). noul: open -68.07; sealed -85.64 (chance-corrected competence). score: open 41.40; sealed 37.36 (chance-corrected competence). | 0.1395% CI 0.01–0.49 | 4.76 | 81.71 | 80.28 | 48.40 | -0.44 | 0.05249estimate |
| 84 ≈ | jeffLogan Markewich · jev-rebuildEvaluation setupSource entry: jeff (Logan Markewich, GLiFormer 400M). Published attribution: Logan Markewich. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 7.137 s; p95 37.170 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.08; C 0.08. choice: open 33.01; sealed 21.79 (chance-corrected competence). noul: open -28.09; sealed -60.76 (chance-corrected competence). score: open 33.21; sealed 28.29 (chance-corrected competence). | 0.1295% CI 0.00–0.51 | 4.43 | 80.27 | 55.76 | 81.09 | -3.56 | 0.00427estimate |
| 85 | smalljev semantic-v9Aditya (isHeSatoshi) · jev-rebuildEvaluation setupSource entry: smalljev semantic-v9. Published attribution: Aditya (isHeSatoshi). evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.451 s; p95 0.501 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.05; C 0.05. choice: open 31.14; sealed 18.27 (chance-corrected competence). noul: open -40.40; sealed -63.82 (chance-corrected competence). score: open 42.06; sealed 34.99 (chance-corrected competence). | 0.0795% CI 0.00–0.38 | 3.66 | 73.09 | 86.46 | 60.77 | -3.52 | 0.02030estimate |
| 86 | Laya multilingualConvai Innovations · system-one-openEvaluation setupSource entry: Laya multilingual. Published attribution: Convai Innovations. evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1) Adjusted latency: p50 1.036 s; p95 4.097 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor. Secondary composites: B 0.01; C 0.01. choice: open 20.06; sealed 9.13 (chance-corrected competence). noul: open -35.98; sealed -31.05 (chance-corrected competence). score: open 24.29; sealed 27.68 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 0.0295% CI 0.00–0.27 | 2.35 | 43.58 | 73.72 | 82.16 | 1.92 | 0.00393estimate |
| 87 ≈ | OpenDecision ModernBERT-largeDeepan Wadhwa · classifierEvaluation setupSource entry: OpenDecision (ModernBERT-large zero-shot). Published attribution: Deepan Wadhwa. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.297 s; p95 0.740 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.01; C 0.01. choice: open 34.17; sealed 20.04 (chance-corrected competence). noul: open -47.14; sealed -70.91 (chance-corrected competence). score: open 41.38; sealed 35.08 (chance-corrected competence). | 0.0195% CI 0.00–0.19 | 2.07 | 72.52 | 86.58 | 79.11 | -5.26 | 0.00497estimate |
| 88 ≈ | BAAI bge-reranker-v2-m3BAAI · rerankerEvaluation setupSource entry: BAAI bge-reranker-v2-m3. Published attribution: BAAI. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.209 s; p95 0.422 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: base-model market reference (deepinfra:BAAI/bge-m3) Secondary composites: B 0.00; C 0.00. choice: open 4.41; sealed 4.20 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 32.30; sealed 27.38 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 83.33 | 90.55 | 59.33 | -22.80 | 0.02269estimate |
| 89 ≈ | Certo v1AltSlate Labs · jev-rebuildEvaluation setupSource entry: Certo v1 (AltSlate Labs). Published attribution: AltSlate Labs. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.264 s; p95 0.273 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open -1.85; sealed -2.05 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 29.49; sealed 27.25 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 87.52 | 91.43 | 96.91 | -24.93 | 0.00127estimate |
| 90 ≈ | Decision Fast v53aFlyMy.AI · jev-rebuildEvaluation setupSource entry: Decision Fast (FlyMy.AI, v53a). Published attribution: FlyMy.AI (@denti). evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.266 s; p95 0.270 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 38.36; sealed 18.49 (chance-corrected competence). noul: open -59.07; sealed -86.84 (chance-corrected competence). score: open 42.64; sealed 33.26 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 76.08 | 91.44 | 80.09 | -11.70 | 0.00461estimate |
| 91 ≈ | Alibaba GTE Reranker ModernBERT-baseAlibaba · rerankerEvaluation setupSource entry: Alibaba GTE Reranker ModernBERT-base. Published attribution: Alibaba-NLP. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.212 s; p95 0.336 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: size-class proxy (deepinfra:thenlper/gte-base); no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 13.02; sealed 2.35 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 31.09; sealed 28.18 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 75.20 | 91.49 | 69.27 | -23.16 | 0.01058estimate |
| 92 ≈ | kev 0.5BJared Palmer · jev-rebuildEvaluation setupSource entry: kev 0.5B. Published attribution: Jared Palmer. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.357 s; p95 0.400 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 25.24; sealed 11.19 (chance-corrected competence). noul: open -48.02; sealed -94.55 (chance-corrected competence). score: open 33.12; sealed 16.56 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 64.98 | 88.45 | 80.09 | -22.26 | 0.00461estimate |
| 93 ≈ | LayaConvai Innovations · jev-rebuildEvaluation setupSource entry: Laya (Convai Innovations, ModernBERT-large 421M). Published attribution: Convai Innovations. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 1.485 s; p95 2.747 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 37.10; sealed 12.77 (chance-corrected competence). noul: open -75.74; sealed -81.53 (chance-corrected competence). score: open 37.21; sealed 29.20 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 73.71 | 73.89 | 84.89 | -13.19 | 0.00319estimate |
| 94 ≈ | lev-350mFranck Verrot (franckverrot) · jev-rebuildEvaluation setupSource entry: lev-350m (Franck Verrot, LFM2.5-350M). Published attribution: Franck Verrot (franckverrot). evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.194 s; p95 0.223 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 33.06; sealed 15.01 (chance-corrected competence). noul: open -52.20; sealed -100.00 (chance-corrected competence). score: open 35.72; sealed 30.94 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 77.68 | 93.63 | 80.04 | -18.01 | 0.00463estimate |
| 95 ≈ | Qwen3.5-0.8B Decision ModelMourad Ghafiri · jev-rebuildEvaluation setupSource entry: Qwen3.5-0.8B Decision Model (Mourad Ghafiri). Published attribution: Mourad Ghafiri. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 1.308 s; p95 5.023 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 0.00; C 0.00. choice: open 34.93; sealed 40.44 (chance-corrected competence). noul: open -69.95; sealed -86.04 (chance-corrected competence). score: open 45.21; sealed 34.46 (chance-corrected competence). | 0.0095% CI 0.00–0.02 | 0.00 | 73.57 | 71.82 | 79.52 | -3.71 | 0.00482estimate |
| 96 ≈ | Mixedbread mxbai-rerank-base-v2Mixedbread · rerankerEvaluation setupSource entry: Mixedbread mxbai-rerank-base-v2. Published attribution: Mixedbread. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.235 s; p95 0.495 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-0.6B); no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 5.84; sealed 0.62 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 31.19; sealed 27.26 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 86.61 | 89.34 | 60.34 | -24.04 | 0.02100estimate |
| 97 ≈ | Needle 3 (2-bit)Cactus Compute · small-tool-modelEvaluation setupSource entry: Needle 3 (Cactus, 2-bit, local CPU). Published attribution: Cactus Compute. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 135.513 s; p95 285.427 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open -3.57; sealed -8.14 (chance-corrected competence). noul: open -14.52; sealed -23.60 (chance-corrected competence). score: open -15.26; sealed -31.60 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 0.00 | 34.13 | 61.57 | -21.11 | 0.01910estimate |
| 98 ≈ | Needle 3 (options as tools)Cactus Compute · small-tool-modelEvaluation setupSource entry: Needle 3, options as tools (post-hoc adapter mode). Published attribution: Cactus Compute. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 58.160 s; p95 136.690 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 10.87; sealed 2.20 (chance-corrected competence). noul: open -12.38; sealed -30.11 (chance-corrected competence). score: open -15.85; sealed -31.79 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 0.00 | 41.00 | 61.57 | -19.90 | 0.01910estimate |
| 99 ≈ | open-jev-deberta-v3-largeKotoba Labs · jev-rebuildEvaluation setupSource entry: open-jev-deberta-v3-large (local CPU). Published attribution: Kotoba Labs. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 2.933 s; p95 5.097 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 21.00; sealed 16.66 (chance-corrected competence). noul: open -79.16; sealed -87.38 (chance-corrected competence). score: open 30.24; sealed 26.89 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 77.06 | 68.25 | 77.59 | -14.61 | 0.00558estimate |
| 100 ≈ | Open Jev JSON CanvasJoshuaSP · jev-rebuildEvaluation setupSource entry: Open Jev JSON Canvas (JoshuaSP). Published attribution: JoshuaSP. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.437 s; p95 0.635 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 81.05; sealed 76.36 (chance-corrected competence). noul: open 74.92; sealed 74.00 (chance-corrected competence). score: open 81.49; sealed 74.99 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 77.14 | 0.00 | 85.57 | 49.30 | 75.12 | 0.04900estimate |
| 101 ≈ | openJev VerdictHemant (heman10x) · jev-rebuildEvaluation setupSource entry: openJev Verdict (heman10x, ModernBERT-base 151M). Published attribution: Hemant (heman10x). evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.402 s; p95 1.018 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 26.78; sealed 16.64 (chance-corrected competence). noul: open -15.15; sealed -79.16 (chance-corrected competence). score: open 26.09; sealed 20.26 (chance-corrected competence). | 0.0095% CI 0.00–0.01 | 0.00 | 52.22 | 83.88 | 86.62 | -14.09 | 0.00279estimate |
| 102 ≈ | openJev Verdict 1.4Hemant (heman10x) · jev-rebuildEvaluation setupSource entry: openJev Verdict 1.4. Published attribution: Hemant (heman10x). evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.771 s; p95 1.120 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 24.29; sealed 17.86 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 31.47; sealed 27.30 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 80.33 | 80.64 | 86.62 | -18.28 | 0.00279estimate |
| 103 ≈ | Qwen3.8-27B (Chutes TEE)Alibaba · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: Qwen3.8 27B (Chutes TEE). Published attribution: Qwen / Chutes. the operator's hosted API Adjusted latency: p50 6.489 s; p95 31.856 s. none (hosted API, measured as is) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 0.00; C 0.00. choice: open 97.18; sealed 97.66 (chance-corrected competence). noul: open 86.68; sealed 98.80 (chance-corrected competence). score: open 94.83; sealed 98.22 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 95.56 | 98.09 | 56.85 | 0.00 | 98.23 | 2.17839estimate |
| 104 ≈ | verdict-smallManavarya09 (Manav) · jev-rebuildEvaluation setupSource entry: verdict-small (Manavarya09, multilingual-e5-small 118M). Published attribution: Manavarya09 (Manav). evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.218 s; p95 2.854 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 26.45; sealed 12.71 (chance-corrected competence). noul: open -55.15; sealed -43.89 (chance-corrected competence). score: open 25.85; sealed 22.34 (chance-corrected competence). | 0.0095% CI 0.00–0.01 | 0.00 | 59.45 | 82.06 | 100.00 | -2.95 | 0.00087estimate |
| 105 | Von 395Mwfzyx (Victor Hugo) · jev-rebuildEvaluation setupSource entry: Von (wfzyx, Option-Marker 395M). Published attribution: wfzyx (Victor Hugo). evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.922 s; p95 2.906 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 34.42; sealed 17.86 (chance-corrected competence). noul: open -78.58; sealed -98.00 (chance-corrected competence). score: open 38.49; sealed 30.48 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 83.48 | 75.72 | 82.71 | -16.55 | 0.00377estimate |
| 106 | Laya typed-decisionsConvai Innovations · system-one-openEvaluation setupSource entry: Laya typed-decisions. Published attribution: Convai Innovations. evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1) Adjusted latency: p50 3.317 s; p95 13.840 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor. Secondary composites: B 0.00; C 0.00. choice: open 31.18; sealed 17.84 (chance-corrected competence). noul: open -91.90; sealed -99.20 (chance-corrected competence). score: open 35.34; sealed 30.59 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 0.0095% CI 0.00–0.00 | 0.00 | 83.34 | 63.38 | 82.53 | -16.92 | 0.00382estimate |
| Unranked | classifier.dev (fast)classifier.dev · jev-serviceAPI: sealed item text sent to the operatorruns on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2)Evaluation setupSource entry: classifier.dev (fast tier). Published attribution: mrmps (@michael_chomsky). the operator's hosted API Adjusted latency: p50 0.518 s; p95 1.221 s. none (hosted API, measured as is) ESTIMATE (proxy tokens, I-2): operator list price: higher of 19 Sep plan cost and 26 Sep usage tariff USD 0.042/M input (rule 1.2); no exact base-model floor applies Secondary composites: B 74.88; C 74.66. choice: open 84.17; sealed 85.11 (chance-corrected competence). noul: open 56.78; sealed 64.44 (chance-corrected competence). score: open 82.50; sealed 81.87 (chance-corrected competence). | 74.6695% CI 73.50–75.20 | 75.81 | 89.19 | 81.99 | 58.89 | 77.14 | 0.02345estimate |
| Unranked | SimpleJev Qwen3.6-35B-A3BFeatherless AI · jev-rebuildAPI: sealed item text sent to the operatorPartial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked.Evaluation setupSource entry: SimpleJev Qwen3.6-35B-A3B. Published attribution: Featherless AI. the author's public demo endpoint Adjusted latency: p50 1.659 s; p95 1.827 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 0.00; C 0.00. choice: open 16.22; sealed 8.94 (chance-corrected competence). noul: open -36.29; sealed -42.69 (chance-corrected competence). score: open -32.01; sealed -35.15 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 78.50 | 75.18 | 35.15 | -22.97 | 0.14512estimate |
| Unranked | Decision-4B (Eval Engine / Chromia)Eval Engine / Chromia · system-one-openPARTIAL / UNRANKED: 1,550 of 1,624 valid answers; 74 context overflows in our evaluator at 2,048 tokens. No official score or rank.Evaluation setupSource entry: Decision-4B (Eval Engine / Chromia). Published attribution: Eval Engine / Chromia. evaluator-owned H100 GPU Adjusted latency: p50 0.233 s; p95 0.379 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B. Secondary composites: B Not measured; C Not measured. v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | 0.01372estimate |
| Unranked | Jobe Qwen3.5-4BMantisShrimpdev · Not reportedNot measured in JevBench v1.5; no current score or rank.Evaluation setupSource entry: Jobe Qwen3.5-4B (frozen). Published attribution: MantisShrimpdev. Endpoint setup not reported. Adjusted latency: p50 Not measured s; p95 Not measured s. Cost basis not reported. | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not measuredUnreported basis |
| Unranked | Mica v0.1 4Bsky7350 · Not reportedThe frozen refusal policy stopped the run after 1,088 of 1,624 rows: 27 documented refusals were mapped to HTTP 422, then three consecutive passthrough HTTP 400 refusals triggered exit 6. The remaining 536 rows have no scores, so this system is not eligible for an official rank.Evaluation setupSource entry: mica-v01-4b. Published attribution: unknown. Endpoint setup not reported. Adjusted latency: p50 Not measured s; p95 Not measured s. Cost basis not reported. v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not measuredUnreported basis |
| Unranked | OpenJev DiffusionGemma 26B-A4B NVFP4razorback16 / Codiv · Not reportedNot measured in JevBench v1.5; no current score or rank.Evaluation setupSource entry: OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16). Published attribution: razorback16 / Codiv. Endpoint setup not reported. Adjusted latency: p50 Not measured s; p95 Not measured s. Cost basis not reported. | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not measuredUnreported basis |
How to read this leaderboard
The current official option A combines Intelligence, Calibration, Speed, and Cost with equal weights in a harmonic mean, with gates and a generalization penalty. Choice, Noul, and Score each count one third. Intelligence weights open and sealed decisions equally. A statistical tie means the published paired 95% difference interval contains zero.
Operator receipt: 112 sourced rows are currently displayable on this page; the published table names Cygnet and Winnow-12B Q8 as joint leaders (statistical tie).
Honest limit: Florian selected the equal-axis, equal-type headline after reviewing results, as disclosed in the method amendment. Latency includes assumed adjustments, and many costs use frozen base-model estimates. New roster additions have individual intervals without newly inferred pairwise ties. This composite stays outside overall and category scoring; 1.4 and 1.5 are separate protocols.
How to read the JevBench results
JevBench is Benchmark Heaven's evaluation, created by Florian Standhartinger and contributors. Each row names the evaluated model or system and its developer. Reasoning settings, adapters, and serving details describe the configuration that produced the result.
We mirror the complete v1.5.4 roster: 106 ranked systems, one honorable mention, two partial entries, and three incomplete or unmeasured entries. Official ranks and numbers stay attached to the exact configuration. An absent score stays missing; a measured zero stays zero.
The current headline is option A: an equal-weight harmonic mean of Intelligence, Calibration, Speed, and Cost. Choice, Noul, and Score each receive one third. Intelligence gives equal weight to 904 open and 720 sealed decisions. The original frozen method named option B as its headline; Florian selected option A after reviewing results, as disclosed in the headline amendment.
We retain published 95% composite intervals and adjacent-pair statistical ties. Cygnet and Winnow-12B Q8 are joint leaders in the official table. Addendum systems have individual intervals, but missing pairwise markers establish neither a tie nor a separation. The sealed Intelligence column is chance-corrected competence, not percent accuracy.
Cost is USD per 1,000 decisions, with the source tariff or estimate and frozen pricing assumptions retained. It does not establish a token tariff for the evaluated configuration. This composite stays outside overall and category scores. Version 1.4 used different tasks and scoring, so movement between versions is not a like-for-like comparison.
Snapshot
How JevBench changed
This page follows the latest imported release, currently v1.5.4. Earlier results remain labeled by version on model pages. Changes to the tasks and scoring prevent a direct comparison of scores across the 1.4 and 1.5 protocols.
| Release | Decisions per complete run | What changed |
|---|---|---|
| v1.5.4 · current imported release | 904 open + 720 sealed | The roster has 112 systems and 106 official ranks. Manchego v2.1 joins after a complete run; earlier scores, intervals, and the scoring method remain unchanged. |
| v1.5 · protocol change | 904 open + 720 sealed | Choice, Noul, and Score receive equal weight, and sealed results carry half of Intelligence. The official headline changed to option A after the method owner reviewed results. Published intervals and paired statistical ties help distinguish small gaps. |
| v1.4.2.2 · September 27, 2026 | 534 frozen + 308 sealed | Imajev-4B joined a roster of 95 systems, with 91 ranked. Version 1.4 used a harmonic composite and a 20% sealed share of Intelligence; later roster additions preserved earlier measurements. |
Cygnet (73.70) and Winnow-12B Q8 (73.23) are joint leaders in the published table (statistical tie). The table retains individual score intervals and published comparisons between adjacent rows.
112 systems, 106 ranked have been evaluated on JevBench. The benchmark falls in the Decision Models category. JevBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About JevBench
Year
2026
Tasks
904 open decisions plus 720 sealed decisions per complete run
Format
Official option A harmonic composite with published uncertainty
Difficulty
Choice, Noul, and Score decisions with sealed generalization tests
We mirror the complete v1.5.4 roster, including six unranked entries, and preserve every published option A score, rank, confidence interval, and adjacent-pair tie. The table keeps configurations separate and displays missing measurements explicitly. The original frozen method and its later headline amendment remain linked beside the data.
Freshness and provenance
Version
JevBench v1.5.4
Refresh cadence
Pinned release
Staleness state
Current
Question availability
601 of 904 open decisions published; 720 sealed decisions private
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does JevBench 1.5 measure?
JevBench 1.5 evaluates typed decisions using Choice, Noul, and Score requests across 904 open and 720 sealed items. Its current official option A combines Intelligence, Calibration, Speed, and Cost equally. The published table includes composite confidence intervals and statistical ties; its score is an index rather than percent accuracy.
Why are some JevBench 1.5 systems unranked?
The v1.5.4 roster contains 112 systems and 106 official ranks. classifier.dev remains an honorable mention because it runs another entrant. Two entries have partial evaluations, and three are incomplete or unmeasured. We retain their published evidence and reasons while leaving unavailable scores and ranks missing.
Are JevBench 1.4 and 1.5 scores comparable?
The versions use different task sets and scoring. Version 1.5 doubles the sealed share of Intelligence to 50%, adds native evaluation of three decision types, and publishes uncertainty. Model evidence rows retain their versioned benchmark keys, and this page explains the earlier releases. A higher score on 1.5 does not by itself establish that a model improved.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.