WinoGrande
Resolve an ambiguous reference in a sentence.
Decision models take application state, a written question, and an allowed answer space. Their output fits a check your code can act on. A bounded answer can still be wrong, and reported confidence needs testing on the decisions your application actually makes.
| Primitive | Example task | Output |
|---|---|---|
| Choice | Route a support ticket to billing, technical support, or account support. | An option and its probability distribution. |
| Noul | Check whether a request meets a written approval rule. | A probability that a yes/no proposition holds. |
| Score | Rate a bug against an ordered severity rubric. | A score and probabilities over the allowed levels. |
The OpenRouter Jev explainer shows each primitive with a real API response.
Perplexity released Decider v1 27B on October 1, 2026, as a Qwen3.8-27B fine-tune with Apache 2.0 weights. It reads text, JSON, and images, then returns yes/no probabilities, choices, or rubric scores. The hosted Decisions API charges $0.04 per million input tokens, with free output and no per-request fee.
Fastino released GLiDE on September 30, 2026. It returns Noul, Choice, and Score answers with probabilities and confidence, and spends additional reasoning on uncertain decisions. The API accepts 40,000 tokens per rendered question prompt. Fastino lists $0.30 per million input tokens and free output; input usage sums internal passes across questions.
Fastino reports 64.81 skill points for GLiDE on Decision Index 0.2.1, compared with 57.91 for the published Jev 1.13.0 reference. Its September 30, 2026 announcement uses the official scorer; GLiDE is not listed on the public board.
| Metric | Unit | GLiDE | TypeSafe Jev 1.13.0 |
|---|---|---|---|
| Decision Index 0.2.1 | Skill points | 64.81 | 57.91 |
| Knowledge and Reasoning | Skill points | 62.9 | 51.4 |
| Language Understanding | Skill points | 63.8 | 62.0 |
| Retrieval and Classification | Skill points | 60.9 | 55.4 |
| Tools and Automation | Skill points | 83.5 | 75.1 |
| Arts and Human Taste | Skill points | 46.0 | 37.7 |
| CLadder | Accuracy | 88.7% | 72.6% |
| CRUXEval | Accuracy | 92.6% | 73.0% |
| Evaluated system | Skill points | Source qualification |
|---|---|---|
| d1 | 58.9 | Liquid AI self-reported reproduction, quoted by Fastino |
| Surogate Rune 26B-A4B v3 | 57.44 | Published Decision Index reference quoted by Fastino |
| Decider chat · Gemma-4-31B | 57.33 | Inference technique on Gemma-4-31B; published reference quoted by Fastino |
| AutoJev-27B | 56.40 | Published Decision Index reference quoted by Fastino |
| simple-jev · Qwen3.8-27B | 55.74 | Featherless inference technique; published reference quoted by Fastino |
The five area scores and overall index are chance-adjusted skill points. CLadder and CRUXEval are raw accuracies. Fastino reports a complete run, but publishes exact GLiDE values for only these eight metrics. Radar-chart gaps do not supply the other individual benchmark accuracies. We have not rerun the evaluation, and these results stay outside general model rankings. The launch chart also quotes five other overall references, including Liquid AI’s self-reported d1 reproduction; their category and task scores are not inferred.
Decision Index report and limits · Fastino launch and charts
Perplexity reports 85.71% accuracy for Decider v1 27B, 84.51% for TypeSafe Jev 1.13.0, and 74.76% for Qwen3.8-27B on a fixed 7,210-row panel. The chart covers September 2026 and was published October 1. Decider was measured through the Perplexity API.
These are provider-reported results. The task samples have different sizes, so the published overall weights them differently. Prompts, sampled item identities, label mappings, baseline inference settings, and uncertainty are not specified in the model card. This panel stays separate from full-dataset results and general model rankings; JevBench public-hard accuracy is separate from its full composite.
| Benchmark | Rows | Jev 1.13.0 | Qwen3.8-27B | Perplexity Decider |
|---|---|---|---|---|
| WinoGrande | 1,000 | 90.70% | 73.10% | 83.30% |
| FinancialPhraseBank | 999 | 76.98% | 75.68% | 84.18% |
| RAGTruth | 1,500 | 77.27% | 61.53% | 88.80% |
| JudgeBench | 350 | 78.57% | 68.86% | 78.29% |
| BBH | 750 | 94.27% | 72.80% | 82.80% |
| JevBench public hard | 101 | 73.27% | 72.28% | 70.30% |
| TabFact | 500 | 89.80% | 78.60% | 90.60% |
| ContractNLI | 510 | 77.45% | 80.78% | 80.78% |
| Circa | 500 | 84.60% | 87.00% | 89.20% |
| Belebele | 500 | 95.00% | 93.20% | 94.00% |
| TruthfulQA binary | 500 | 92.00% | 82.80% | 85.40% |
| Overall | 7,210 | 84.51% | 74.76% | 85.71% |
Emphasized values mark the best percentage in each row, including ties. Read the panel result notes, the pinned Perplexity model card, or the official launch and chart.
JevBench v1.5.4 covers 1624 decisions. Its official score combines Intelligence, Calibration, Speed, and Cost; it is an index, rather than percent accuracy. The 112 configurations include 106 ranked systems and retain the source's unranked entries.
The headline uses option A, selected after the method owner reviewed the results. Published composite intervals and adjacent-pair statistical ties remain attached to each configuration. Cost per 1,000 decisions retains source tariffs or estimates. This composite stays outside general model rankings.
Option A · joint leaders (statistical tie). The ≈ marker identifies a published statistical tie with the next row. Missing pairwise markers establish neither a tie nor a separation; 95% intervals below describe individual composite scores.
| Rank | System | Score | Intelligence | Calibration | Speed | Cost | Sealed Intelligence | USD / 1,000 decisions |
|---|---|---|---|---|---|---|---|---|
| 1 ≈ | Cygnetblockbrain · system-one-openEvaluation setupSource entry: Cygnet (blockbrain, frozen Gemma-4-12B-it). Published attribution: blockbrain. evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.230 s; p95 0.348 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 73.16; C 73.70. choice: open 82.92; sealed 80.29 (chance-corrected competence). noul: open 57.86; sealed 55.67 (chance-corrected competence). score: open 76.60; sealed 73.20 (chance-corrected competence). | 73.7095% CI 72.36–74.46 | 71.09 | 87.01 | 90.97 | 56.43 | 69.72 | 0.02834estimate |
| 2 | Winnow-12B Q8Eldan Ring · jev-rebuildEvaluation setupSource entry: Winnow-12B Q8. Published attribution: Eldan Ring. evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.340 s; p95 0.715 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 73.47; C 73.23. choice: open 77.61; sealed 82.04 (chance-corrected competence). noul: open 58.61; sealed 66.15 (chance-corrected competence). score: open 81.21; sealed 80.99 (chance-corrected competence). | 73.2395% CI 72.02–73.99 | 74.43 | 84.07 | 86.14 | 56.56 | 76.39 | 0.02806estimate |
| 3 | Jev 1.13.0TypeSafe AI · jevAPI: sealed item text sent to the operatorEvaluation setupSource entry: Jev 1.13.0 (TypeSafe AI). Published attribution: TypeSafe AI. the operator's hosted API Adjusted latency: p50 0.616 s; p95 0.674 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1); no exact base-model floor applies Secondary composites: B 72.11; C 72.13. choice: open 85.73; sealed 87.59 (chance-corrected competence). noul: open 47.75; sealed 48.62 (chance-corrected competence). score: open 81.16; sealed 81.14 (chance-corrected competence). | 72.1395% CI 71.01–72.61 | 72.00 | 88.03 | 83.81 | 54.73 | 72.45 | 0.03230estimate |
| 4 | JevK5 v0.3 4Ballebee · unclassifiedEvaluation setupSource entry: JevK5 v0.3 (4B). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.182 s; p95 0.243 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 68.11; C 63.24. choice: open 80.25; sealed 74.25 (chance-corrected competence). noul: open 42.33; sealed 14.98 (chance-corrected competence). score: open 62.22; sealed 63.59 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 71.9095% CI 69.39–72.95 | 56.27 | 88.34 | 93.55 | 63.07 | 50.94 | 0.01702estimate |
| 5 | Plumb-4Bcrh225 · unclassifiedEvaluation setupSource entry: Plumb-4B (crh225, JevK5 v0.2 + LoRA). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.185 s; p95 0.245 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 67.75; C 62.00. choice: open 80.43; sealed 73.30 (chance-corrected competence). noul: open 42.95; sealed 16.58 (chance-corrected competence). score: open 59.26; sealed 62.57 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 71.5695% CI 69.20–72.73 | 55.85 | 87.44 | 93.45 | 63.07 | 50.81 | 0.01702estimate |
| 6 ≈ | Jev-Omniakhilaaa3 · jev-rebuildEvaluation setupSource entry: Jev-Omni (akhilaaa3, Gemma-4-12B merged). Published attribution: akhilaaa3. evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.375 s; p95 0.908 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 71.30; C 71.50. choice: open 83.37; sealed 78.70 (chance-corrected competence). noul: open 55.28; sealed 56.73 (chance-corrected competence). score: open 72.79; sealed 76.07 (chance-corrected competence). | 71.5095% CI 70.21–72.40 | 70.49 | 82.60 | 84.68 | 56.05 | 70.50 | 0.02917estimate |
| 7 | decider-4b v2Mapika · system-one-openEvaluation setupSource entry: decider-4b v2 (Mapika). Published attribution: Mapika. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.203 s; p95 0.405 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 67.53; C 61.59. choice: open 80.43; sealed 77.31 (chance-corrected competence). noul: open 42.92; sealed 18.73 (chance-corrected competence). score: open 61.88; sealed 53.36 (chance-corrected competence). | 71.2895% CI 69.11–72.35 | 55.77 | 85.59 | 90.86 | 64.54 | 49.80 | 0.01520estimate |
| 8 | Decision 4B v1.2FlyMy.AI · unclassifiedEvaluation setupSource entry: Decision 4B v1.2 (FlyMyJev, Qwen3.5-4B + LoRA). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.183 s; p95 0.244 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 66.57; C 56.66. choice: open 80.87; sealed 73.76 (chance-corrected competence). noul: open 20.13; sealed 16.44 (chance-corrected competence). score: open 67.79; sealed 63.00 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 70.8395% CI 68.52–71.97 | 53.66 | 88.56 | 93.50 | 63.07 | 51.06 | 0.01702estimate |
| 9 | Imajev-4B (RTX 5090)mohit67890 · unclassifiedEvaluation setupSource entry: Imajev-4B (RTX 5090). Published attribution: unknown. evaluator-owned Lium GPU pod (RTX 5090), offline read-only container Adjusted latency: p50 0.234 s; p95 0.329 s. x2 + 0.15 s (assumption, not measured) ESTIMATE (I-2): measured input tokens; zero generated output tokens for signed logits readout. M2 floor uses the 25 Sep 2026 DeepInfra Qwen3.5-4B snapshot rates (USD 0.03/M input, USD 0.15/M output); the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen3.5-9B. Secondary composites: B 66.20; C 55.91. choice: open 80.12; sealed 82.33 (chance-corrected competence). noul: open 21.82; sealed 7.56 (chance-corrected competence). score: open 66.80; sealed 62.21 (chance-corrected competence). v1.5 roster addendum A2. No pairwise comparison is inferred for this addition. | 70.3995% CI 67.80–71.61 | 53.47 | 88.12 | 91.13 | 63.26 | 50.70 | 0.01677estimate |
| 10 | Decision 4B v1.1FlyMy.AI · unclassifiedEvaluation setupSource entry: Decision 4B v1.1 (FlyMyJev, Qwen3.5-4B + LoRA). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.183 s; p95 0.243 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 66.09; C 55.14. choice: open 80.12; sealed 74.77 (chance-corrected competence). noul: open 18.02; sealed 15.75 (chance-corrected competence). score: open 65.82; sealed 64.17 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 70.3995% CI 66.86–71.57 | 53.11 | 87.31 | 93.52 | 63.07 | 51.56 | 0.01702estimate |
| 11 | Manchego v2.1oraculumai · system-one-openEvaluation setupSource entry: Manchego v2.1. Published attribution: oraculumai. evaluator-owned offline RTX6000 GPU Adjusted latency: p50 0.319 s; p95 0.430 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: frozen 25 Sep DeepInfra Qwen/Qwen3.5-4B market reference, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B). $0.03/M input, zero generated output; 836,500 measured input tokens across 1,624 decisions. Frozen v1.5 base-model floor applied; no bookable Manchego tariff claimed. Re-score if the basis changes. Secondary composites: B 64.38; C 50.11. choice: open 66.94; sealed 65.04 (chance-corrected competence). noul: open 34.26; sealed 16.84 (chance-corrected competence). score: open 62.03; sealed 62.10 (chance-corrected competence). v1.5 roster addendum A5. No pairwise comparison is inferred for this addition. | 68.8195% CI 59.64–70.28 | 51.20 | 84.93 | 88.63 | 64.33 | 47.99 | 0.01545estimate |
| 12 ≈ | SemIf Qwen3.5-4BTheodore Lee (TheoLeeCJ) · jev-rebuildEvaluation setupSource entry: SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ). Published attribution: Theodore Lee (TheoLeeCJ). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.229 s; p95 0.356 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 64.31; C 50.22. choice: open 73.80; sealed 70.96 (chance-corrected competence). noul: open 30.30; sealed 13.93 (chance-corrected competence). score: open 57.89; sealed 61.00 (chance-corrected competence). | 68.6695% CI 60.06–70.16 | 51.31 | 83.97 | 90.89 | 63.07 | 48.63 | 0.01702estimate |
| 13 ≈ | spark-s1-4b-v6Abhishek Rai (abhishek085) · jev-rebuildEvaluation setupSource entry: spark-s1-4b-v6 (Open Spark Jev, abhishek085). Published attribution: Abhishek Rai (abhishek085). evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.406 s; p95 0.690 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 66.86; C 68.16. choice: open 72.55; sealed 73.11 (chance-corrected competence). noul: open 50.44; sealed 46.47 (chance-corrected competence). score: open 64.39; sealed 65.74 (chance-corrected competence). | 68.1695% CI 66.27–69.70 | 62.12 | 69.71 | 85.53 | 60.42 | 61.77 | 0.02086estimate |
| 14 ≈ | metask-jev-4bWayfind (metask-ai) · jev-rebuildEvaluation setupSource entry: metask-jev-4b. Published attribution: Wayfind (metask-ai). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.285 s; p95 0.394 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 64.12; C 53.61. choice: open 70.11; sealed 72.41 (chance-corrected competence). noul: open 43.23; sealed 25.20 (chance-corrected competence). score: open 56.88; sealed 53.04 (chance-corrected competence). | 67.4895% CI 65.39–68.69 | 53.48 | 82.74 | 89.49 | 57.73 | 50.22 | 0.02564estimate |
| 15 ≈ | HopperHopitAI · jev-rebuildEvaluation setupSource entry: Hopper. Published attribution: HopitAI. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.395 s; p95 0.479 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 62.93; C 46.85. choice: open 70.12; sealed 75.48 (chance-corrected competence). noul: open 10.23; sealed 12.76 (chance-corrected competence). score: open 65.76; sealed 64.82 (chance-corrected competence). | 67.4795% CI 56.34–69.15 | 49.86 | 87.87 | 87.23 | 62.26 | 51.02 | 0.01812estimate |
| 16 | Malkuth-4Bnewfull5 (dhtocks) · jev-rebuildEvaluation setupSource entry: Malkuth-4B (newfull5, Kev post-train). Published attribution: newfull5 (dhtocks). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.442 s; p95 0.548 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 63.88; C 54.99. choice: open 71.17; sealed 69.55 (chance-corrected competence). noul: open 39.48; sealed 22.84 (chance-corrected competence). score: open 64.44; sealed 59.23 (chance-corrected competence). | 66.7795% CI 64.94–67.89 | 54.45 | 83.21 | 86.16 | 55.82 | 50.54 | 0.02968estimate |
| 17 | Surogate Rune 26B-A4B v3 (RTX PRO 6000)Surogate · unclassifiedEvaluation setupSource entry: Surogate Rune 26B-A4B v3 (RTX PRO 6000). Published attribution: unknown. evaluator-owned Lium GPU pod (RTX PRO 6000), offline read-only container Adjusted latency: p50 0.351 s; p95 0.721 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens. Secondary composites: B 66.55; C 66.47. choice: open 81.99; sealed 82.17 (chance-corrected competence). noul: open 45.59; sealed 49.53 (chance-corrected competence). score: open 79.28; sealed 79.88 (chance-corrected competence). v1.5 roster addendum A2. No pairwise comparison is inferred for this addition. | 66.4795% CI 65.43–67.01 | 69.74 | 88.30 | 85.97 | 48.97 | 70.53 | 0.05024estimate |
| 18 ≈ | reflex 4Bkshetrajna12 · jev-rebuildEvaluation setupSource entry: reflex 4B (kshetrajna12). Published attribution: kshetrajna12. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 2.865 s; p95 4.627 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 61.93; C 48.05. choice: open 73.66; sealed 69.31 (chance-corrected competence). noul: open 33.33; sealed 19.45 (chance-corrected competence). score: open 58.62; sealed 54.56 (chance-corrected competence). | 65.2495% CI 58.56–66.35 | 51.49 | 86.84 | 68.77 | 63.14 | 47.78 | 0.01693estimate |
| 19 ≈ | jev-local Qwen3.5-9Bus (GitHub) · jev-rebuildEvaluation setupSource entry: jev-local (Qwen3.5-9B). Published attribution: us (GitHub). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.944 s; p95 5.339 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 63.22; C 57.39. choice: open 66.88; sealed 65.39 (chance-corrected competence). noul: open 46.94; sealed 37.02 (chance-corrected competence). score: open 60.71; sealed 60.71 (chance-corrected competence). | 65.2495% CI 63.46–66.50 | 56.28 | 77.82 | 72.98 | 58.86 | 54.37 | 0.02352estimate |
| 20 ≈ | djevMaisa · jev-rebuildEvaluation setupSource entry: djev (Maisa, diffusion-gemma). Published attribution: Maisa (David Villalón). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.251 s; p95 0.317 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 64.76; C 64.16. choice: open 76.62; sealed 72.78 (chance-corrected competence). noul: open 66.16; sealed 70.25 (chance-corrected competence). score: open 75.33; sealed 72.80 (chance-corrected competence). | 64.1695% CI 63.07–65.00 | 72.33 | 80.41 | 91.00 | 48.22 | 71.95 | 0.05321estimate |
| 21 ≈ | Raw Qwen3 4B Instruct 2507 direct logitsAlibaba · raw-logit-controlEvaluation setupSource entry: Raw Qwen3 4B Instruct 2507 direct logits. Published attribution: Alibaba Qwen / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.262 s; p95 0.456 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 60.33; C 50.45. choice: open 65.50; sealed 63.46 (chance-corrected competence). noul: open 49.81; sealed 38.04 (chance-corrected competence). score: open 56.98; sealed 50.60 (chance-corrected competence). | 62.1495% CI 59.35–64.22 | 54.06 | 52.95 | 89.23 | 63.36 | 50.70 | 0.01665estimate |
| 22 ≈ | jqv Qwen3-32BOctalab · jev-rebuildEvaluation setupSource entry: jqv (Qwen3-32B zero-shot). Published attribution: hjmurmur (Octalab). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.307 s; p95 1.503 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 57.50; C 42.20. choice: open 73.00; sealed 73.09 (chance-corrected competence). noul: open 18.67; sealed 10.18 (chance-corrected competence). score: open 62.35; sealed 57.23 (chance-corrected competence). | 60.7795% CI 51.25–64.26 | 49.09 | 86.77 | 83.36 | 51.17 | 46.83 | 0.04242estimate |
| 23 ≈ | JevK5 v0.2.0allebee · jev-rebuildEvaluation setupSource entry: JevK5 v0.2.0. Published attribution: allebee. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.216 s; p95 0.379 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 53.55; C 40.36. choice: open 77.61; sealed 71.76 (chance-corrected competence). noul: open 12.19; sealed -3.89 (chance-corrected competence). score: open 60.23; sealed 62.33 (chance-corrected competence). | 58.1195% CI 47.56–68.24 | 46.70 | 84.87 | 90.86 | 63.07 | 43.40 | 0.01702estimate |
| 24 | Qwen3.5-9B Jev-like data-mix v2jsaurabh · jev-rebuildEvaluation setupSource entry: Qwen3.5-9B Jev-like data-mix v2. Published attribution: jsaurabh. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.581 s; p95 1.087 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 52.49; C 53.02. choice: open 72.91; sealed 74.77 (chance-corrected competence). noul: open 42.37; sealed 31.67 (chance-corrected competence). score: open 72.56; sealed 68.41 (chance-corrected competence). | 53.0295% CI 51.88–53.88 | 60.45 | 80.86 | 82.00 | 45.69 | 58.28 | 0.06461estimate |
| 25 ≈ | Standard One 8BStandard Thinking · jev-rebuildEvaluation setupSource entry: Standard One 8B (Standard Thinking). Published attribution: Standard Thinking (myeongho12). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.203 s; p95 0.278 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 47.13; C 47.15. choice: open 73.72; sealed 72.18 (chance-corrected competence). noul: open 45.04; sealed 42.98 (chance-corrected competence). score: open 61.11; sealed 62.56 (chance-corrected competence). | 47.7995% CI 46.67–48.54 | 59.60 | 83.16 | 92.48 | 43.28 | 59.24 | 0.07772estimate |
| 26 | NInfer Qwen3.8-Flash-Next mixedIgor L. / NInfer contributors · native-logitEvaluation setupSource entry: NInfer Qwen3.8-Flash-Next mixed. Published attribution: Igor L. / NInfer contributors. evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.310 s; p95 0.445 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 47.70; C 47.47. choice: open 84.93; sealed 79.11 (chance-corrected competence). noul: open 51.40; sealed 43.02 (chance-corrected competence). score: open 74.44; sealed 70.30 (chance-corrected competence). | 47.4795% CI 46.70–47.88 | 67.20 | 88.54 | 88.61 | 42.53 | 64.14 | 0.08233estimate |
| 27 | Instinct Dual 4BZooWork · decision-apiAPI: sealed item text sent to the operatorEvaluation setupSource entry: Instinct Dual 4B. Published attribution: ZooWork / pierre-srp. operator-hosted free-preview API Adjusted latency: p50 0.506 s; p95 1.260 s. x2 (demo assumption, not measured) ESTIMATE: ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule); reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B. Secondary composites: B 42.97; C 32.62. choice: open 75.05; sealed 74.49 (chance-corrected competence). noul: open 4.41; sealed -12.80 (chance-corrected competence). score: open 59.41; sealed 58.15 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 46.9795% CI 37.86–56.77 | 43.12 | 88.30 | 81.95 | 60.19 | 39.95 | 0.02124estimate |
| 28 ≈ | swanOneblockbrain · system-one-openEvaluation setupSource entry: swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4). Published attribution: blockbrain. evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.568 s; p95 0.591 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 47.32; C 46.56. choice: open 88.11; sealed 82.65 (chance-corrected competence). noul: open 45.55; sealed 59.89 (chance-corrected competence). score: open 77.84; sealed 73.31 (chance-corrected competence). | 46.5695% CI 45.81–46.94 | 71.22 | 87.06 | 84.74 | 42.15 | 71.95 | 0.08478estimate |
| 29 ≈ | Raw Qwen3 8B direct logitsAlibaba · raw-logit-controlEvaluation setupSource entry: Raw Qwen3 8B direct logits. Published attribution: Alibaba Qwen / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.333 s; p95 0.621 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 44.63; C 32.85. choice: open 64.86; sealed 56.55 (chance-corrected competence). noul: open 48.07; sealed 38.40 (chance-corrected competence). score: open 57.40; sealed 41.56 (chance-corrected competence). | 45.2295% CI 36.89–46.74 | 51.14 | 49.19 | 86.84 | 45.54 | 45.50 | 0.06539estimate |
| 30 ≈ | decider-2bMapika · jev-rebuildEvaluation setupSource entry: decider-2b (Mapika). Published attribution: Mapika. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.177 s; p95 0.205 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 41.11; C 31.32. choice: open 52.02; sealed 53.48 (chance-corrected competence). noul: open 35.24; sealed 5.78 (chance-corrected competence). score: open 59.11; sealed 48.39 (chance-corrected competence). | 45.1095% CI 33.82–54.26 | 42.34 | 71.53 | 94.40 | 64.92 | 35.89 | 0.01477estimate |
| 31 ≈ | system-one Qwen3-8BSean Goedecke · jev-rebuildEvaluation setupSource entry: system-one (Qwen3-8B, Sean Goedecke). Published attribution: Sean Goedecke. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.227 s; p95 0.360 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 43.43; C 31.23. choice: open 63.18; sealed 56.68 (chance-corrected competence). noul: open 49.89; sealed 40.04 (chance-corrected competence). score: open 47.55; sealed 45.51 (chance-corrected competence). | 44.1495% CI 36.79–45.66 | 50.47 | 49.38 | 90.88 | 44.97 | 47.41 | 0.06829estimate |
| 32 ≈ | system-one-openmithalouni · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: system-one-open (Gemma 4 E2B LoRA on an L4). Published attribution: mithalouni. the author's public demo endpoint Adjusted latency: p50 1.150 s; p95 1.295 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 38.80; C 29.47. choice: open 59.92; sealed 52.04 (chance-corrected competence). noul: open 15.40; sealed 14.98 (chance-corrected competence). score: open 57.99; sealed 49.52 (chance-corrected competence). | 42.4495% CI 33.84–51.92 | 41.64 | 71.63 | 78.27 | 68.33 | 38.85 | 0.01137estimate |
| 33 ≈ | Autoloops Gemma 4 31B ITAutoloops · jevAPI: sealed item text sent to the operatorEvaluation setupSource entry: Autoloops – Gemma 4 31B IT. Published attribution: Autoloops. the operator's hosted API Adjusted latency: p50 0.607 s; p95 0.667 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1) Secondary composites: B 41.84; C 40.53. choice: open 86.36; sealed 87.51 (chance-corrected competence). noul: open 68.73; sealed 67.45 (chance-corrected competence). score: open 75.55; sealed 74.51 (chance-corrected competence). | 40.5395% CI 40.09–40.76 | 76.68 | 85.84 | 83.92 | 39.59 | 76.49 | 0.10323tariff |
| 34 | GPT-6 Luna (low)OpenAI · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: GPT-6 Luna (low reasoning effort). Published attribution: OpenAI. the operator's hosted API Adjusted latency: p50 1.580 s; p95 3.031 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1); no exact base-model floor applies Secondary composites: B 43.10; C 40.48. choice: open 99.06; sealed 97.87 (chance-corrected competence). noul: open 85.25; sealed 90.18 (chance-corrected competence). score: open 99.87; sealed 99.52 (chance-corrected competence). | 40.4895% CI 40.27–40.66 | 95.29 | 94.92 | 73.20 | 39.06 | 95.86 | 0.10750estimate |
| 35 ≈ | GPT-6 Luna (medium)OpenAI · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: GPT-6 Luna (default medium reasoning effort). Published attribution: OpenAI. the operator's hosted API Adjusted latency: p50 1.555 s; p95 3.088 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1); no exact base-model floor applies Secondary composites: B 41.35; C 38.75. choice: open 99.38; sealed 99.51 (chance-corrected competence). noul: open 90.69; sealed 89.93 (chance-corrected competence). score: open 98.89; sealed 99.02 (chance-corrected competence). | 38.7595% CI 38.57–38.93 | 96.24 | 95.56 | 73.19 | 38.32 | 96.15 | 0.11379estimate |
| 36 ≈ | JevOneJuspay · jev-rebuildEvaluation setupSource entry: JevOne. Published attribution: Juspay. evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.267 s; p95 0.339 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 37.40; C 31.13. choice: open 82.43; sealed 77.14 (chance-corrected competence). noul: open 9.38; sealed 17.13 (chance-corrected competence). score: open 69.91; sealed 68.81 (chance-corrected competence). | 38.2595% CI 37.38–38.79 | 54.13 | 84.94 | 90.43 | 39.84 | 54.36 | 0.10122estimate |
| 37 ≈ | kev 4BJared Palmer · jev-rebuildEvaluation setupSource entry: kev 4B (research preview). Published attribution: Jared Palmer. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.493 s; p95 0.574 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 34.59; C 26.44. choice: open 53.89; sealed 46.76 (chance-corrected competence). noul: open 29.06; sealed 4.29 (chance-corrected competence). score: open 58.24; sealed 48.55 (chance-corrected competence). | 38.0795% CI 27.48–46.77 | 39.86 | 67.64 | 85.48 | 65.78 | 33.20 | 0.01382estimate |
| 38 ≈ | kev 8BJared Palmer · jev-rebuildEvaluation setupSource entry: kev 8B (research preview). Published attribution: Jared Palmer. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.511 s; p95 0.754 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 33.09; C 23.72. choice: open 60.67; sealed 58.72 (chance-corrected competence). noul: open 38.20; sealed 14.00 (chance-corrected competence). score: open 64.04; sealed 54.14 (chance-corrected competence). | 34.1595% CI 26.79–37.48 | 48.30 | 71.02 | 84.14 | 40.42 | 42.28 | 0.09682estimate |
| 39 ≈ | open-alternative-jev Qwen3.5-4BIkerMoel · jev-rebuildEvaluation setupSource entry: open-alternative-jev (Qwen3.5-4B, IkerMoel). Published attribution: IkerMoel. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.223 s; p95 0.334 s. x2 + 0.15 s (assumption, not measured) ESTIMATE (proxy tokens, I-2): exact base-model market reference; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 29.92; C 23.31. choice: open 63.28; sealed 65.00 (chance-corrected competence). noul: open 5.29; sealed -16.76 (chance-corrected competence). score: open 59.22; sealed 48.12 (chance-corrected competence). | 33.5695% CI 25.75–41.66 | 37.36 | 76.87 | 91.29 | 63.28 | 32.12 | 0.01675estimate |
| 40 ≈ | Bespoke Nimble 9BBespoke Labs · jev-rebuildEvaluation setupSource entry: Bespoke Nimble 9B (Bespoke Labs). Published attribution: Bespoke Labs. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.547 s; p95 0.927 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 32.31; C 31.82. choice: open 71.68; sealed 72.82 (chance-corrected competence). noul: open 52.86; sealed 44.55 (chance-corrected competence). score: open 70.90; sealed 69.68 (chance-corrected competence). | 31.8295% CI 31.19–32.27 | 63.75 | 77.16 | 82.95 | 36.75 | 62.35 | 0.12833estimate |
| 41 ≈ | Malkuth-2Bnewfull5 (dhtocks) · jev-rebuildEvaluation setupSource entry: Malkuth-2B (newfull5, Kev post-train). Published attribution: newfull5 (dhtocks). evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.231 s; p95 0.286 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 26.38; C 20.76. choice: open 54.83; sealed 48.49 (chance-corrected competence). noul: open 26.01; sealed -5.78 (chance-corrected competence). score: open 51.90; sealed 43.12 (chance-corrected competence). | 29.9095% CI 20.96–38.98 | 35.54 | 75.19 | 91.80 | 65.57 | 28.61 | 0.01405estimate |
| 42 | openjev-sglang Qwen3.6-35B-A3Bekzhang · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: openjev-sglang (Qwen3.6-35B-A3B on SGLang). Published attribution: ekzhang. the author's public demo endpoint Adjusted latency: p50 1.196 s; p95 1.305 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 29.18; C 27.73. choice: open 80.55; sealed 74.60 (chance-corrected competence). noul: open 36.38; sealed 27.20 (chance-corrected competence). score: open 67.14; sealed 66.02 (chance-corrected competence). | 29.0395% CI 28.42–29.40 | 58.65 | 82.87 | 78.07 | 35.63 | 55.94 | 0.13981estimate |
| 43 ≈ | decider-35b-a3bMapika · jev-rebuildEvaluation setupSource entry: decider-35b-a3b (Mapika). Published attribution: Mapika. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.241 s; p95 0.330 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 27.71; C 27.50. choice: open 74.30; sealed 76.47 (chance-corrected competence). noul: open 51.34; sealed 31.31 (chance-corrected competence). score: open 66.79; sealed 62.55 (chance-corrected competence). | 27.5095% CI 26.93–27.87 | 60.46 | 82.04 | 91.00 | 34.39 | 56.78 | 0.15387estimate |
| 44 | local-jev Qwen3.5-4BAmith Chandrappa (amithgc) · jev-rebuildEvaluation setupSource entry: local-jev Qwen3.5-4B. Published attribution: Amith Chandrappa (amithgc). evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.427 s; p95 0.991 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 22.73; C 17.93. choice: open 73.67; sealed 69.06 (chance-corrected competence). noul: open -23.24; sealed -38.29 (chance-corrected competence). score: open 59.68; sealed 61.61 (chance-corrected competence). | 25.8295% CI 19.89–32.96 | 33.75 | 82.45 | 83.74 | 59.27 | 30.79 | 0.02279estimate |
| 45 | Nemotron Diffusion 8B (optimized vLLM)pst2154 · system-one-openEvaluation setupSource entry: Nemotron Diffusion 8B (pst2154, optimized vLLM). Published attribution: pst2154. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.187 s; p95 0.225 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 22.70; C 17.84. choice: open 57.05; sealed 58.81 (chance-corrected competence). noul: open 10.77; sealed -16.91 (chance-corrected competence). score: open 50.79; sealed 42.55 (chance-corrected competence). v1.5 roster addendum A3. No pairwise comparison is inferred for this addition. | 25.6895% CI 18.60–32.78 | 33.84 | 76.25 | 93.74 | 55.49 | 28.15 | 0.03045estimate |
| 46 ≈ | Open-Jev 9BZefan Cai (@Zefan_Cai) · jev-rebuildEvaluation setupSource entry: Open-Jev 9B (Zefan Cai). Published attribution: Zefan Cai (@Zefan_Cai). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 1.276 s; p95 3.432 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 24.99; C 24.36. choice: open 66.19; sealed 71.45 (chance-corrected competence). noul: open 48.96; sealed 58.69 (chance-corrected competence). score: open 70.21; sealed 67.42 (chance-corrected competence). | 24.3695% CI 23.93–24.64 | 63.82 | 81.53 | 73.59 | 33.05 | 65.86 | 0.17041estimate |
| 47 ≈ | Decision 2B v59FlyMy.AI · jev-rebuildEvaluation setupSource entry: Decision 2B (FlyMy.AI, v59). Published attribution: FlyMy.AI (@denti). evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.300 s; p95 0.307 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 19.30; C 15.64. choice: open 58.27; sealed 67.70 (chance-corrected competence). noul: open -16.11; sealed -35.13 (chance-corrected competence). score: open 61.78; sealed 51.36 (chance-corrected competence). | 22.5295% CI 16.68–28.77 | 31.31 | 86.22 | 90.36 | 66.37 | 27.98 | 0.01321estimate |
| 48 ≈ | GPT-5.6 Luna (low)OpenAI · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: GPT-5.6 Luna (low reasoning effort). Published attribution: OpenAI. the operator's hosted API Adjusted latency: p50 1.321 s; p95 2.970 s. none (hosted API, measured as is) operator list price Secondary composites: B 24.16; C 22.37. choice: open 95.94; sealed 93.85 (chance-corrected competence). noul: open 87.11; sealed 94.15 (chance-corrected competence). score: open 96.42; sealed 98.52 (chance-corrected competence). | 22.3795% CI 22.26–22.46 | 94.33 | 94.71 | 74.06 | 30.67 | 95.51 | 0.20467tariff |
| 49 ≈ | typecastlmMikhail Gribov · system-one-openEvaluation setupSource entry: typecastlm (Mikhail Gribov, Qwen3.5-4B computed head). Published attribution: Mikhail Gribov. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.225 s; p95 0.278 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 18.85; C 15.16. choice: open 66.45; sealed 72.01 (chance-corrected competence). noul: open -52.09; sealed -7.49 (chance-corrected competence). score: open 56.73; sealed 51.83 (chance-corrected competence). | 21.8395% CI 16.23–28.58 | 31.24 | 76.71 | 92.04 | 63.96 | 38.78 | 0.01590estimate |
| 50 ≈ | JEV Qwen3.5-9B Base NVFP4WilfLin · jev-rebuildEvaluation setupSource entry: JEV Qwen3.5-9B Base NVFP4. Published attribution: WilfLin. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.190 s; p95 0.234 s. x2 + 0.15 s (assumption, not measured) ESTIMATE (proxy tokens, I-2): exact base-model market reference Secondary composites: B 17.76; C 13.94. choice: open 71.17; sealed 59.87 (chance-corrected competence). noul: open -34.85; sealed -22.47 (chance-corrected competence). score: open 62.32; sealed 57.43 (chance-corrected competence). | 20.0795% CI 15.19–25.51 | 32.24 | 80.93 | 93.54 | 47.59 | 31.61 | 0.05584estimate |
| 51 | Gemini 3.1 Flash-LiteGoogle · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: Gemini 3.1 Flash-Lite. Published attribution: Google. the operator's hosted API Adjusted latency: p50 0.863 s; p95 1.144 s. none (hosted API, measured as is) operator list price Secondary composites: B 20.78; C 19.58. choice: open 83.98; sealed 84.09 (chance-corrected competence). noul: open 73.61; sealed 75.05 (chance-corrected competence). score: open 72.25; sealed 76.53 (chance-corrected competence). | 19.5895% CI 19.30–19.83 | 77.59 | 74.68 | 80.05 | 29.76 | 78.56 | 0.21941tariff |
| 52 | AutoJev-27Bdenis-pplx · unclassifiedEvaluation setupSource entry: AutoJev-27B (denis-pplx, Qwen3.8-27B). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.352 s; p95 0.492 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 20.45; C 19.54. choice: open 85.24; sealed 82.24 (chance-corrected competence). noul: open 50.44; sealed 62.00 (chance-corrected competence). score: open 80.06; sealed 76.70 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 19.5495% CI 19.30–19.63 | 72.78 | 87.70 | 87.62 | 29.37 | 73.65 | 0.22619estimate |
| 53 | AutoJev-27B (RTX PRO 6000)denis-pplx · unclassifiedEvaluation setupSource entry: AutoJev-27B (RTX PRO 6000). Published attribution: unknown. evaluator-owned Lium GPU pod (RTX PRO 6000), offline read-only container Adjusted latency: p50 0.326 s; p95 0.602 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: M2 base-model market reference from the frozen 25 Sep 2026 OpenRouter snapshot applied to measured tokens. Secondary composites: B 20.39; C 19.48. choice: open 84.93; sealed 81.70 (chance-corrected competence). noul: open 51.14; sealed 62.00 (chance-corrected competence). score: open 79.98; sealed 76.76 (chance-corrected competence). v1.5 roster addendum A2. No pairwise comparison is inferred for this addition. | 19.4895% CI 19.26–19.60 | 72.75 | 86.68 | 87.07 | 29.37 | 73.49 | 0.22619estimate |
| 54 | NInfer Qwen3.8-27B NVFP4Igor L. / NInfer contributors · native-logitEvaluation setupSource entry: NInfer Qwen3.8-27B NVFP4. Published attribution: Igor L. / NInfer contributors. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.248 s; p95 0.416 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 19.35; C 18.74. choice: open 80.86; sealed 80.00 (chance-corrected competence). noul: open 47.77; sealed 44.73 (chance-corrected competence). score: open 72.29; sealed 67.19 (chance-corrected competence). | 18.7495% CI 18.45–18.92 | 65.47 | 85.91 | 89.86 | 29.12 | 63.97 | 0.23052estimate |
| 55 | Eikos-27Bcaiovicentino1 · unclassifiedEvaluation setupSource entry: Eikos-27B (caiovicentino1, Qwen3.8-27B). Published attribution: unknown. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.354 s; p95 0.488 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 19.48; C 18.50. choice: open 87.62; sealed 88.27 (chance-corrected competence). noul: open 56.47; sealed 64.15 (chance-corrected competence). score: open 78.35; sealed 75.66 (chance-corrected competence). v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | 18.5095% CI 18.28–18.60 | 75.08 | 86.29 | 87.63 | 28.69 | 76.02 | 0.23824estimate |
| 56 ≈ | NInfer Qwen3.8-27B NVFP4 (T=1.5)Igor L. / NInfer contributors · native-logitEvaluation setupSource entry: NInfer Qwen3.8-27B NVFP4 (T=1.5). Published attribution: Igor L. / NInfer contributors. evaluator-owned Lium GPU pod (RTX5090), offline read-only container Adjusted latency: p50 0.248 s; p95 0.416 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 18.89; C 18.48. choice: open 80.86; sealed 80.00 (chance-corrected competence). noul: open 36.83; sealed 36.36 (chance-corrected competence). score: open 69.61; sealed 62.80 (chance-corrected competence). | 18.4895% CI 18.18–18.65 | 61.08 | 86.49 | 89.86 | 29.12 | 59.72 | 0.23052estimate |
| 57 | Instinct Qwen3.8-27BZooWork · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: Instinct (ZooWork, Qwen3.8-27B). Published attribution: rayrain-srp (ZooWork). the author's public demo endpoint Adjusted latency: p50 0.519 s; p95 1.308 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 18.86; C 18.32. choice: open 84.49; sealed 79.67 (chance-corrected competence). noul: open 28.50; sealed 39.56 (chance-corrected competence). score: open 72.75; sealed 71.52 (chance-corrected competence). | 18.3295% CI 18.03–18.51 | 62.75 | 84.96 | 81.68 | 29.16 | 63.58 | 0.22979estimate |
| 58 | OpenJev (thinking, BF16)razorback16 · jev-rebuildEvaluation setupSource entry: OpenJev (thinking, BF16). Published attribution: razorback16. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 1.561 s; p95 2.632 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 19.27; C 17.94. choice: open 81.67; sealed 78.51 (chance-corrected competence). noul: open 80.61; sealed 88.62 (chance-corrected competence). score: open 85.26; sealed 90.57 (chance-corrected competence). | 17.9495% CI 17.76–18.08 | 84.21 | 83.10 | 73.86 | 28.51 | 85.90 | 0.24146estimate |
| 59 | djev (thinking)Maisa · jev-rebuildEvaluation setupSource entry: djev (thinking). Published attribution: David Villalon / Maisa. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 1.481 s; p95 3.959 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 18.47; C 17.40. choice: open 64.77; sealed 48.29 (chance-corrected competence). noul: open 74.63; sealed 95.75 (chance-corrected competence). score: open 90.27; sealed 90.39 (chance-corrected competence). | 17.4095% CI 17.28–17.51 | 77.35 | 95.66 | 72.32 | 28.13 | 78.14 | 0.24866estimate |
| 60 | LitJev Qwen3.8-27BZhengxu Yu · jev-rebuildEvaluation setupSource entry: LitJev (Qwen3.8-27B). Published attribution: Zhengxu Yu. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 3.068 s; p95 4.906 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 16.74; C 15.41. choice: open 80.44; sealed 79.03 (chance-corrected competence). noul: open 23.49; sealed 33.71 (chance-corrected competence). score: open 70.26; sealed 63.02 (chance-corrected competence). | 16.3195% CI 16.03–16.48 | 58.32 | 84.45 | 68.22 | 28.36 | 58.59 | 0.24438estimate |
| 61 | Bev / Bonsai 27BReza Sayar · system-one-openEvaluation setupSource entry: Bev / Bonsai 27B. Published attribution: Reza Sayar. evaluator-owned H100 GPU Adjusted latency: p50 2.191 s; p95 2.433 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate. Frozen market reference qwen/qwen3.8-27b. Secondary composites: B 15.99; C 12.34. choice: open 74.11; sealed 76.27 (chance-corrected competence). noul: open 4.29; sealed 29.49 (chance-corrected competence). score: open 67.71; sealed 66.61 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 15.7795% CI 15.13–16.05 | 53.08 | 77.68 | 72.73 | 28.23 | 57.46 | 0.24676estimate |
| 62 ≈ | Raw Phi-4 mini direct logitsMicrosoft · raw-logit-controlEvaluation setupSource entry: Raw Phi-4 mini direct logits. Published attribution: Microsoft / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.281 s; p95 0.431 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 13.11; C 10.57. choice: open 54.10; sealed 56.67 (chance-corrected competence). noul: open 6.03; sealed -43.64 (chance-corrected competence). score: open 54.69; sealed 46.85 (chance-corrected competence). | 15.2295% CI 10.26–20.74 | 27.62 | 71.41 | 89.17 | 53.22 | 19.96 | 0.03625estimate |
| 63 ≈ | OpenSourceJev Qwen3.5-4B Q4_K_Msabeel111 · jev-rebuildEvaluation setupSource entry: OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp). Published attribution: sabeel111. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 1.389 s; p95 4.041 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3) Secondary composites: B 11.16; C 9.20. choice: open 64.38; sealed 51.49 (chance-corrected competence). noul: open -36.12; sealed -37.24 (chance-corrected competence). score: open 55.31; sealed 56.79 (chance-corrected competence). | 13.2595% CI 8.98–18.31 | 25.77 | 76.04 | 72.51 | 69.14 | 23.68 | 0.01069estimate |
| 64 | reflex-27bkshetrajna12 · jev-rebuildEvaluation setupSource entry: reflex-27b (Qwen3.8-27B). Published attribution: kshetrajna12. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 2.748 s; p95 4.379 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 13.77; C 13.19. choice: open 84.80; sealed 80.30 (chance-corrected competence). noul: open 30.92; sealed 35.71 (chance-corrected competence). score: open 73.55; sealed 71.49 (chance-corrected competence). | 13.1995% CI 12.99–13.30 | 62.79 | 85.93 | 69.20 | 25.80 | 62.50 | 0.29734estimate |
| 65 ≈ | Open-Jev 2BZefan Cai (@Zefan_Cai) · jev-rebuildEvaluation setupSource entry: Open-Jev 2B (Zefan Cai). Published attribution: Zefan Cai (@Zefan_Cai). evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 1.008 s; p95 2.631 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 8.44; C 6.30. choice: open 51.52; sealed 45.49 (chance-corrected competence). noul: open 17.34; sealed -8.15 (chance-corrected competence). score: open 47.50; sealed 47.69 (chance-corrected competence). | 9.0795% CI 6.62–11.54 | 33.56 | 73.62 | 75.76 | 33.05 | 28.34 | 0.17041estimate |
| 66 ≈ | GLiNER2 largeFastino · classifierEvaluation setupSource entry: GLiNER2 large (Fastino). Published attribution: Fastino. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 1.879 s; p95 17.719 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 7.19; C 5.83. choice: open 26.57; sealed 28.42 (chance-corrected competence). noul: open 16.35; sealed -0.62 (chance-corrected competence). score: open 35.16; sealed 29.07 (chance-corrected competence). | 8.4095% CI 5.21–12.34 | 22.49 | 42.38 | 64.78 | 77.59 | 18.96 | 0.00558estimate |
| 67 ≈ | Qwen3-Reranker-4BAlibaba · rerankerEvaluation setupSource entry: Qwen3-Reranker-4B. Published attribution: Qwen. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.511 s; p95 2.140 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: hosted exact-model reference (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies Secondary composites: B 5.86; C 4.90. choice: open 56.70; sealed 45.60 (chance-corrected competence). noul: open -10.06; sealed -45.93 (chance-corrected competence). score: open 44.40; sealed 40.42 (chance-corrected competence). | 7.0695% CI 4.26–10.74 | 21.03 | 76.26 | 79.61 | 48.40 | 13.36 | 0.05249estimate |
| 68 ≈ | DeepSeek V4.1 FlashDeepSeek · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: DeepSeek V4.1 Flash (thinking default). Published attribution: DeepSeek. the operator's hosted API Adjusted latency: p50 1.776 s; p95 6.465 s. none (hosted API, measured as is) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 7.41; C 6.65. choice: open 96.20; sealed 98.23 (chance-corrected competence). noul: open 86.17; sealed 96.55 (chance-corrected competence). score: open 93.16; sealed 91.82 (chance-corrected competence). | 6.6595% CI 6.62–6.66 | 93.69 | 96.92 | 69.40 | 19.09 | 95.53 | 0.49760estimate |
| 69 ≈ | SimpleJev Qwen3.5-0.8BFeatherless AI · jev-rebuildEvaluation setupSource entry: SimpleJev (Qwen3.5-0.8B, CPU). Published attribution: sabeel111 / Featherless AI. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 7.371 s; p95 16.793 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 3.26; C 2.77. choice: open 28.28; sealed 27.77 (chance-corrected competence). noul: open -12.83; sealed 4.25 (chance-corrected competence). score: open 24.54; sealed 28.47 (chance-corrected competence). | 4.0095% CI 2.06–6.66 | 16.75 | 46.67 | 59.07 | 70.31 | 20.16 | 0.00976estimate |
| 70 ≈ | SimpleJev Qwen3.8-27BFeatherless AI · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: SimpleJev Qwen3.8-27B. Published attribution: Featherless AI. the author's public demo endpoint Adjusted latency: p50 1.679 s; p95 1.920 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 3.71; C 3.36. choice: open 85.43; sealed 80.67 (chance-corrected competence). noul: open 51.10; sealed 60.95 (chance-corrected competence). score: open 79.20; sealed 79.18 (chance-corrected competence). | 3.3695% CI 3.33–3.37 | 72.75 | 87.20 | 74.92 | 14.90 | 73.60 | 0.68680estimate |
| 71 ≈ | decision-machine-1milliseconds.ai · decision-apiAPI: sealed item text sent to the operatorEvaluation setupSource entry: decision-machine-1 (milliseconds.ai). Published attribution: milliseconds.ai (Baptiste Laget). the operator's hosted API Adjusted latency: p50 0.180 s; p95 0.293 s. none (hosted API, measured as is) operator standard launch list price (interpretation I-1); no exact base-model floor applies Secondary composites: B 2.50; C 2.25. choice: open 53.27; sealed 44.94 (chance-corrected competence). noul: open -17.55; sealed -67.96 (chance-corrected competence). score: open 46.78; sealed 38.41 (chance-corrected competence). | 3.2495% CI 1.73–5.39 | 14.82 | 81.18 | 92.77 | 56.30 | 5.13 | 0.02863estimate |
| 72 | GLiNER2.5 multiFastino · classifierEvaluation setupSource entry: GLiNER2.5 multi (Fastino, 287M). Published attribution: Fastino. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 1.370 s; p95 16.069 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 2.10; C 1.89. choice: open 19.44; sealed 17.24 (chance-corrected competence). noul: open 0.16; sealed -2.58 (chance-corrected competence). score: open 27.74; sealed 21.98 (chance-corrected competence). | 2.7295% CI 1.20–5.10 | 14.00 | 57.96 | 66.57 | 86.62 | 12.21 | 0.00279estimate |
| 73 | Bosun v3.1 0.6BClause Logic · system-one-openEvaluation setupSource entry: Bosun v3.1 0.6B. Published attribution: Clause Logic. evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1, local_files_only=True), credential-free (no API key, no HF token) Adjusted latency: p50 4.037 s; p95 16.515 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-base frozen raw-qwen3-0.6b estimate. Secondary composites: B 1.92; C 1.74. choice: open 33.44; sealed 36.69 (chance-corrected competence). noul: open -4.15; sealed -57.85 (chance-corrected competence). score: open 40.17; sealed 37.15 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 2.5095% CI 1.15–4.56 | 13.58 | 65.16 | 61.76 | 77.55 | 5.33 | 0.00560estimate |
| 74 | GLiNER2.5 baseFastino · classifierEvaluation setupSource entry: GLiNER2 (Fastino, gliner2.5-base). Published attribution: Fastino. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.998 s; p95 9.285 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 1.80; C 1.58. choice: open 24.80; sealed 20.99 (chance-corrected competence). noul: open 1.50; sealed -10.00 (chance-corrected competence). score: open 25.42; sealed 18.14 (chance-corrected competence). | 2.2795% CI 0.90–4.50 | 13.48 | 35.57 | 70.33 | 86.62 | 9.71 | 0.00279estimate |
| 75 | Deem 0.8B v1LibertAI · system-one-openEvaluation setupSource entry: Deem 0.8B v1. Published attribution: LibertAI. evaluator-owned H100 GPU Adjusted latency: p50 0.501 s; p95 0.739 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate. Frozen market reference deepinfra:Qwen/Qwen3.5-0.8B. Secondary composites: B 1.66; C 1.48. choice: open 28.48; sealed 19.54 (chance-corrected competence). noul: open 7.67; sealed -28.73 (chance-corrected competence). score: open 37.90; sealed 19.95 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 2.1495% CI 0.90–4.34 | 13.02 | 38.88 | 84.32 | 80.29 | 3.59 | 0.00454estimate |
| 76 ≈ | JevActeinptein · jev-rebuildAPI: sealed item text sent to the operatorEvaluation setupSource entry: JevAct (einptein, jev1-2b-v2). Published attribution: einptein. the author's public demo endpoint Adjusted latency: p50 0.802 s; p95 2.836 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 1.14; C 1.06. choice: open 31.06; sealed 35.28 (chance-corrected competence). noul: open -10.65; sealed -55.35 (chance-corrected competence). score: open 38.74; sealed 30.47 (chance-corrected competence). | 1.5295% CI 0.52–3.18 | 11.24 | 62.54 | 76.43 | 68.56 | 3.47 | 0.01117estimate |
| 77 ≈ | CLM-8BContrastive-LM · system-one-openEvaluation setupSource entry: CLM-8B (Contrastive-LM, clm-latest). Published attribution: Contrastive-LM (Kwok, Kang, Suresh, Saad-Falcon, Pavone, Ré, Mirhoseini). evaluator-owned Lium GPU pod (RTXPRO6000), offline read-only container Adjusted latency: p50 0.183 s; p95 0.256 s. x2 + 0.15 s (assumption, not measured) price floor: base-model reference price applied Secondary composites: B 1.16; C 1.05. choice: open 19.36; sealed 8.29 (chance-corrected competence). noul: open 6.95; sealed -13.49 (chance-corrected competence). score: open 24.96; sealed 22.52 (chance-corrected competence). | 1.5195% CI 0.48–3.17 | 11.43 | 48.48 | 93.29 | 50.50 | 5.77 | 0.04466estimate |
| 78 ≈ | kev 0.6BJared Palmer · jev-rebuildEvaluation setupSource entry: kev 0.6B (research preview). Published attribution: Jared Palmer. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.422 s; p95 0.450 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.92; C 0.88. choice: open 33.25; sealed 21.31 (chance-corrected competence). noul: open 1.96; sealed -57.60 (chance-corrected competence). score: open 41.18; sealed 31.81 (chance-corrected competence). | 1.2695% CI 0.49–2.50 | 10.34 | 67.65 | 87.22 | 80.09 | -1.49 | 0.00461estimate |
| 79 ≈ | Raw Qwen3 0.6B direct logitsAlibaba · raw-logit-controlEvaluation setupSource entry: Raw Qwen3 0.6B direct logits. Published attribution: Alibaba Qwen / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.276 s; p95 0.331 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.86; C 0.75. choice: open 17.12; sealed 23.80 (chance-corrected competence). noul: open 13.42; sealed -4.25 (chance-corrected competence). score: open -0.51; sealed 13.79 (chance-corrected competence). | 1.0895% CI 0.27–2.61 | 10.56 | 21.45 | 90.39 | 77.58 | 11.11 | 0.00559estimate |
| 80 ≈ | GLiNER2.5 smallFastino · classifierEvaluation setupSource entry: GLiNER2.5 small (Fastino, 74M). Published attribution: Fastino. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.472 s; p95 3.891 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.65; C 0.62. choice: open 15.71; sealed 11.87 (chance-corrected competence). noul: open 4.31; sealed -34.11 (chance-corrected competence). score: open 31.47; sealed 27.19 (chance-corrected competence). | 0.8995% CI 0.23–2.20 | 9.19 | 55.90 | 77.36 | 86.62 | 1.65 | 0.00279estimate |
| 81 ≈ | Raw Qwen3 1.7B direct logitsAlibaba · raw-logit-controlEvaluation setupSource entry: Raw Qwen3 1.7B direct logits. Published attribution: Alibaba Qwen / neutral reproduction. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.282 s; p95 0.337 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.67; C 0.59. choice: open 41.04; sealed 23.29 (chance-corrected competence). noul: open 11.66; sealed -3.45 (chance-corrected competence). score: open -15.90; sealed 1.40 (chance-corrected competence). | 0.8595% CI 0.15–2.28 | 9.67 | 21.57 | 90.22 | 68.55 | 7.08 | 0.01118estimate |
| 82 ≈ | MirrorBluusun · jev-rebuildEvaluation setupSource entry: Mirror. Published attribution: Bluusun. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 4.578 s; p95 9.436 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.15; C 0.15. choice: open 1.64; sealed -9.37 (chance-corrected competence). noul: open -0.61; sealed -4.55 (chance-corrected competence). score: open 19.21; sealed 27.26 (chance-corrected competence). | 0.2295% CI 0.01–0.90 | 5.60 | 43.21 | 63.64 | 89.27 | 4.45 | 0.00228estimate |
| 83 ≈ | ZeroEntropy zerank-2ZeroEntropy · rerankerEvaluation setupSource entry: ZeroEntropy zerank-2. Published attribution: ZeroEntropy. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.473 s; p95 1.984 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-4B); no exact base-model floor applies Secondary composites: B 0.09; C 0.09. choice: open 56.57; sealed 46.96 (chance-corrected competence). noul: open -68.07; sealed -85.64 (chance-corrected competence). score: open 41.40; sealed 37.36 (chance-corrected competence). | 0.1395% CI 0.01–0.49 | 4.76 | 81.71 | 80.28 | 48.40 | -0.44 | 0.05249estimate |
| 84 ≈ | jeffLogan Markewich · jev-rebuildEvaluation setupSource entry: jeff (Logan Markewich, GLiFormer 400M). Published attribution: Logan Markewich. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 7.137 s; p95 37.170 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.08; C 0.08. choice: open 33.01; sealed 21.79 (chance-corrected competence). noul: open -28.09; sealed -60.76 (chance-corrected competence). score: open 33.21; sealed 28.29 (chance-corrected competence). | 0.1295% CI 0.00–0.51 | 4.43 | 80.27 | 55.76 | 81.09 | -3.56 | 0.00427estimate |
| 85 | smalljev semantic-v9Aditya (isHeSatoshi) · jev-rebuildEvaluation setupSource entry: smalljev semantic-v9. Published attribution: Aditya (isHeSatoshi). evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.451 s; p95 0.501 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.05; C 0.05. choice: open 31.14; sealed 18.27 (chance-corrected competence). noul: open -40.40; sealed -63.82 (chance-corrected competence). score: open 42.06; sealed 34.99 (chance-corrected competence). | 0.0795% CI 0.00–0.38 | 3.66 | 73.09 | 86.46 | 60.77 | -3.52 | 0.02030estimate |
| 86 | Laya multilingualConvai Innovations · system-one-openEvaluation setupSource entry: Laya multilingual. Published attribution: Convai Innovations. evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1) Adjusted latency: p50 1.036 s; p95 4.097 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor. Secondary composites: B 0.01; C 0.01. choice: open 20.06; sealed 9.13 (chance-corrected competence). noul: open -35.98; sealed -31.05 (chance-corrected competence). score: open 24.29; sealed 27.68 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 0.0295% CI 0.00–0.27 | 2.35 | 43.58 | 73.72 | 82.16 | 1.92 | 0.00393estimate |
| 87 ≈ | OpenDecision ModernBERT-largeDeepan Wadhwa · classifierEvaluation setupSource entry: OpenDecision (ModernBERT-large zero-shot). Published attribution: Deepan Wadhwa. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.297 s; p95 0.740 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.01; C 0.01. choice: open 34.17; sealed 20.04 (chance-corrected competence). noul: open -47.14; sealed -70.91 (chance-corrected competence). score: open 41.38; sealed 35.08 (chance-corrected competence). | 0.0195% CI 0.00–0.19 | 2.07 | 72.52 | 86.58 | 79.11 | -5.26 | 0.00497estimate |
| 88 ≈ | BAAI bge-reranker-v2-m3BAAI · rerankerEvaluation setupSource entry: BAAI bge-reranker-v2-m3. Published attribution: BAAI. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.209 s; p95 0.422 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: base-model market reference (deepinfra:BAAI/bge-m3) Secondary composites: B 0.00; C 0.00. choice: open 4.41; sealed 4.20 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 32.30; sealed 27.38 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 83.33 | 90.55 | 59.33 | -22.80 | 0.02269estimate |
| 89 ≈ | Certo v1AltSlate Labs · jev-rebuildEvaluation setupSource entry: Certo v1 (AltSlate Labs). Published attribution: AltSlate Labs. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.264 s; p95 0.273 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open -1.85; sealed -2.05 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 29.49; sealed 27.25 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 87.52 | 91.43 | 96.91 | -24.93 | 0.00127estimate |
| 90 ≈ | Decision Fast v53aFlyMy.AI · jev-rebuildEvaluation setupSource entry: Decision Fast (FlyMy.AI, v53a). Published attribution: FlyMy.AI (@denti). evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.266 s; p95 0.270 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 38.36; sealed 18.49 (chance-corrected competence). noul: open -59.07; sealed -86.84 (chance-corrected competence). score: open 42.64; sealed 33.26 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 76.08 | 91.44 | 80.09 | -11.70 | 0.00461estimate |
| 91 ≈ | Alibaba GTE Reranker ModernBERT-baseAlibaba · rerankerEvaluation setupSource entry: Alibaba GTE Reranker ModernBERT-base. Published attribution: Alibaba-NLP. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.212 s; p95 0.336 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: size-class proxy (deepinfra:thenlper/gte-base); no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 13.02; sealed 2.35 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 31.09; sealed 28.18 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 75.20 | 91.49 | 69.27 | -23.16 | 0.01058estimate |
| 92 ≈ | kev 0.5BJared Palmer · jev-rebuildEvaluation setupSource entry: kev 0.5B. Published attribution: Jared Palmer. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.357 s; p95 0.400 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 25.24; sealed 11.19 (chance-corrected competence). noul: open -48.02; sealed -94.55 (chance-corrected competence). score: open 33.12; sealed 16.56 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 64.98 | 88.45 | 80.09 | -22.26 | 0.00461estimate |
| 93 ≈ | LayaConvai Innovations · jev-rebuildEvaluation setupSource entry: Laya (Convai Innovations, ModernBERT-large 421M). Published attribution: Convai Innovations. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 1.485 s; p95 2.747 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 37.10; sealed 12.77 (chance-corrected competence). noul: open -75.74; sealed -81.53 (chance-corrected competence). score: open 37.21; sealed 29.20 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 73.71 | 73.89 | 84.89 | -13.19 | 0.00319estimate |
| 94 ≈ | lev-350mFranck Verrot (franckverrot) · jev-rebuildEvaluation setupSource entry: lev-350m (Franck Verrot, LFM2.5-350M). Published attribution: Franck Verrot (franckverrot). evaluator-owned Lium GPU pod (RTX6000), offline read-only container Adjusted latency: p50 0.194 s; p95 0.223 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 33.06; sealed 15.01 (chance-corrected competence). noul: open -52.20; sealed -100.00 (chance-corrected competence). score: open 35.72; sealed 30.94 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 77.68 | 93.63 | 80.04 | -18.01 | 0.00463estimate |
| 95 ≈ | Qwen3.5-0.8B Decision ModelMourad Ghafiri · jev-rebuildEvaluation setupSource entry: Qwen3.5-0.8B Decision Model (Mourad Ghafiri). Published attribution: Mourad Ghafiri. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 1.308 s; p95 5.023 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate Secondary composites: B 0.00; C 0.00. choice: open 34.93; sealed 40.44 (chance-corrected competence). noul: open -69.95; sealed -86.04 (chance-corrected competence). score: open 45.21; sealed 34.46 (chance-corrected competence). | 0.0095% CI 0.00–0.02 | 0.00 | 73.57 | 71.82 | 79.52 | -3.71 | 0.00482estimate |
| 96 ≈ | Mixedbread mxbai-rerank-base-v2Mixedbread · rerankerEvaluation setupSource entry: Mixedbread mxbai-rerank-base-v2. Published attribution: Mixedbread. evaluator-owned Lium GPU pod (A6000), offline read-only container Adjusted latency: p50 0.235 s; p95 0.495 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: size-class proxy (deepinfra:Qwen/Qwen3-Reranker-0.6B); no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 5.84; sealed 0.62 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 31.19; sealed 27.26 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 86.61 | 89.34 | 60.34 | -24.04 | 0.02100estimate |
| 97 ≈ | Needle 3 (2-bit)Cactus Compute · small-tool-modelEvaluation setupSource entry: Needle 3 (Cactus, 2-bit, local CPU). Published attribution: Cactus Compute. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 135.513 s; p95 285.427 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open -3.57; sealed -8.14 (chance-corrected competence). noul: open -14.52; sealed -23.60 (chance-corrected competence). score: open -15.26; sealed -31.60 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 0.00 | 34.13 | 61.57 | -21.11 | 0.01910estimate |
| 98 ≈ | Needle 3 (options as tools)Cactus Compute · small-tool-modelEvaluation setupSource entry: Needle 3, options as tools (post-hoc adapter mode). Published attribution: Cactus Compute. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 58.160 s; p95 136.690 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 10.87; sealed 2.20 (chance-corrected competence). noul: open -12.38; sealed -30.11 (chance-corrected competence). score: open -15.85; sealed -31.79 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 0.00 | 41.00 | 61.57 | -19.90 | 0.01910estimate |
| 99 ≈ | open-jev-deberta-v3-largeKotoba Labs · jev-rebuildEvaluation setupSource entry: open-jev-deberta-v3-large (local CPU). Published attribution: Kotoba Labs. evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 2.933 s; p95 5.097 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 21.00; sealed 16.66 (chance-corrected competence). noul: open -79.16; sealed -87.38 (chance-corrected competence). score: open 30.24; sealed 26.89 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 77.06 | 68.25 | 77.59 | -14.61 | 0.00558estimate |
| 100 ≈ | Open Jev JSON CanvasJoshuaSP · jev-rebuildEvaluation setupSource entry: Open Jev JSON Canvas (JoshuaSP). Published attribution: JoshuaSP. evaluator-owned Lium GPU pod (H100), offline read-only container Adjusted latency: p50 0.437 s; p95 0.635 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 81.05; sealed 76.36 (chance-corrected competence). noul: open 74.92; sealed 74.00 (chance-corrected competence). score: open 81.49; sealed 74.99 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 77.14 | 0.00 | 85.57 | 49.30 | 75.12 | 0.04900estimate |
| 101 ≈ | openJev VerdictHemant (heman10x) · jev-rebuildEvaluation setupSource entry: openJev Verdict (heman10x, ModernBERT-base 151M). Published attribution: Hemant (heman10x). evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.402 s; p95 1.018 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 26.78; sealed 16.64 (chance-corrected competence). noul: open -15.15; sealed -79.16 (chance-corrected competence). score: open 26.09; sealed 20.26 (chance-corrected competence). | 0.0095% CI 0.00–0.01 | 0.00 | 52.22 | 83.88 | 86.62 | -14.09 | 0.00279estimate |
| 102 ≈ | openJev Verdict 1.4Hemant (heman10x) · jev-rebuildEvaluation setupSource entry: openJev Verdict 1.4. Published attribution: Hemant (heman10x). evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.771 s; p95 1.120 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: v1.5 measured-input proxy; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 24.29; sealed 17.86 (chance-corrected competence). noul: open -100.00; sealed -100.00 (chance-corrected competence). score: open 31.47; sealed 27.30 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 80.33 | 80.64 | 86.62 | -18.28 | 0.00279estimate |
| 103 ≈ | Qwen3.8-27B (Chutes TEE)Alibaba · llm-baselineAPI: sealed item text sent to the operatorEvaluation setupSource entry: Qwen3.8 27B (Chutes TEE). Published attribution: Qwen / Chutes. the operator's hosted API Adjusted latency: p50 6.489 s; p95 31.856 s. none (hosted API, measured as is) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 0.00; C 0.00. choice: open 97.18; sealed 97.66 (chance-corrected competence). noul: open 86.68; sealed 98.80 (chance-corrected competence). score: open 94.83; sealed 98.22 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 95.56 | 98.09 | 56.85 | 0.00 | 98.23 | 2.17839estimate |
| 104 ≈ | verdict-smallManavarya09 (Manav) · jev-rebuildEvaluation setupSource entry: verdict-small (Manavarya09, multilingual-e5-small 118M). Published attribution: Manavarya09 (Manav). evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.218 s; p95 2.854 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 26.45; sealed 12.71 (chance-corrected competence). noul: open -55.15; sealed -43.89 (chance-corrected competence). score: open 25.85; sealed 22.34 (chance-corrected competence). | 0.0095% CI 0.00–0.01 | 0.00 | 59.45 | 82.06 | 100.00 | -2.95 | 0.00087estimate |
| 105 | Von 395Mwfzyx (Victor Hugo) · jev-rebuildEvaluation setupSource entry: Von (wfzyx, Option-Marker 395M). Published attribution: wfzyx (Victor Hugo). evaluator-owned CPU container, offline read-only, pinned cores Adjusted latency: p50 0.922 s; p95 2.906 s. x2 + 0.15 s (assumption, not measured) documented hosted-model estimate; no exact base-model floor applies Secondary composites: B 0.00; C 0.00. choice: open 34.42; sealed 17.86 (chance-corrected competence). noul: open -78.58; sealed -98.00 (chance-corrected competence). score: open 38.49; sealed 30.48 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 83.48 | 75.72 | 82.71 | -16.55 | 0.00377estimate |
| 106 | Laya typed-decisionsConvai Innovations · system-one-openEvaluation setupSource entry: Laya typed-decisions. Published attribution: Convai Innovations. evaluator-owned CPU, Sandy (AMD Ryzen 5 3600), 4 of 12 threads, nice -n 5, shared host, offline (HF_HUB_OFFLINE=1) Adjusted latency: p50 3.317 s; p95 13.840 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate; no exact base-model floor applies. $0.01/M input, $0/M output; same-class hosted-encoder estimate; no exact-base market floor. Secondary composites: B 0.00; C 0.00. choice: open 31.18; sealed 17.84 (chance-corrected competence). noul: open -91.90; sealed -99.20 (chance-corrected competence). score: open 35.34; sealed 30.59 (chance-corrected competence). v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | 0.0095% CI 0.00–0.00 | 0.00 | 83.34 | 63.38 | 82.53 | -16.92 | 0.00382estimate |
| Unranked | classifier.dev (fast)classifier.dev · jev-serviceAPI: sealed item text sent to the operatorruns on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2)Evaluation setupSource entry: classifier.dev (fast tier). Published attribution: mrmps (@michael_chomsky). the operator's hosted API Adjusted latency: p50 0.518 s; p95 1.221 s. none (hosted API, measured as is) ESTIMATE (proxy tokens, I-2): operator list price: higher of 19 Sep plan cost and 26 Sep usage tariff USD 0.042/M input (rule 1.2); no exact base-model floor applies Secondary composites: B 74.88; C 74.66. choice: open 84.17; sealed 85.11 (chance-corrected competence). noul: open 56.78; sealed 64.44 (chance-corrected competence). score: open 82.50; sealed 81.87 (chance-corrected competence). | 74.6695% CI 73.50–75.20 | 75.81 | 89.19 | 81.99 | 58.89 | 77.14 | 0.02345estimate |
| Unranked | SimpleJev Qwen3.6-35B-A3BFeatherless AI · jev-rebuildAPI: sealed item text sent to the operatorPartial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked.Evaluation setupSource entry: SimpleJev Qwen3.6-35B-A3B. Published attribution: Featherless AI. the author's public demo endpoint Adjusted latency: p50 1.659 s; p95 1.827 s. x2 demo-endpoint adjustment (assumption, not measured) ESTIMATE: exact base-model market reference (operator tariff excluded by 30-day rule) Secondary composites: B 0.00; C 0.00. choice: open 16.22; sealed 8.94 (chance-corrected competence). noul: open -36.29; sealed -42.69 (chance-corrected competence). score: open -32.01; sealed -35.15 (chance-corrected competence). | 0.0095% CI 0.00–0.00 | 0.00 | 78.50 | 75.18 | 35.15 | -22.97 | 0.14512estimate |
| Unranked | Decision-4B (Eval Engine / Chromia)Eval Engine / Chromia · system-one-openPARTIAL / UNRANKED: 1,550 of 1,624 valid answers; 74 context overflows in our evaluator at 2,048 tokens. No official score or rank.Evaluation setupSource entry: Decision-4B (Eval Engine / Chromia). Published attribution: Eval Engine / Chromia. evaluator-owned H100 GPU Adjusted latency: p50 0.233 s; p95 0.379 s. x2 + 0.15 s (assumption, not measured) ESTIMATE: documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B); re-score if it changes (I-3). Frozen market reference deepinfra:Qwen/Qwen3.5-4B. Secondary composites: B Not measured; C Not measured. v1.5 roster addendum A4. No pairwise comparison is inferred for this addition. | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | 0.01372estimate |
| Unranked | Jobe Qwen3.5-4BMantisShrimpdev · Not reportedNot measured in JevBench v1.5; no current score or rank.Evaluation setupSource entry: Jobe Qwen3.5-4B (frozen). Published attribution: MantisShrimpdev. Endpoint setup not reported. Adjusted latency: p50 Not measured s; p95 Not measured s. Cost basis not reported. | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not measuredUnreported basis |
| Unranked | Mica v0.1 4Bsky7350 · Not reportedThe frozen refusal policy stopped the run after 1,088 of 1,624 rows: 27 documented refusals were mapped to HTTP 422, then three consecutive passthrough HTTP 400 refusals triggered exit 6. The remaining 536 rows have no scores, so this system is not eligible for an official rank.Evaluation setupSource entry: mica-v01-4b. Published attribution: unknown. Endpoint setup not reported. Adjusted latency: p50 Not measured s; p95 Not measured s. Cost basis not reported. v1.5 roster addendum A1. No pairwise comparison is inferred for this addition. | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not measuredUnreported basis |
| Unranked | OpenJev DiffusionGemma 26B-A4B NVFP4razorback16 / Codiv · Not reportedNot measured in JevBench v1.5; no current score or rank.Evaluation setupSource entry: OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16). Published attribution: razorback16 / Codiv. Endpoint setup not reported. Adjusted latency: p50 Not measured s; p95 Not measured s. Cost basis not reported. | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not measuredUnreported basis |
Read the JevBench method and result notes, or inspect the pinned official results.
JevBench by Florian Standhartinger and contributors. MIT license and copyright notice.
DecisionBench by Hanno Labs. A separate benchmark for typed decisions. Its applied suite and reasoning track stay separate; results are linked at the source rather than imported into this table.
The seven original benchmark pages include 40 source tables, with full numeric metrics and profiles for identified evaluated configurations. Each page provides a JSON download. Anonymous submissions and conflicting source values are labeled.
These datasets cover different decisions, from sentiment labels to evidence checks. Compare scores only when the sample, label mapping, language, and evaluation setup match. A subset or binary conversion keeps its own result.
Resolve an ambiguous reference in a sentence.
Classify financial sentiment as positive, negative, or neutral.
Detect unsupported claims against retrieved evidence.
Choose the objectively better response from a pair.
Solve challenging reasoning tasks from BIG-Bench.
Evaluate decision systems; public-subset accuracy is separate from the full composite.
Check whether a table supports or refutes a statement.
Classify a contract hypothesis and identify supporting evidence.
Interpret an indirect answer to a yes/no question.
Answer passage-comprehension questions across language variants.
Test resistance to common misconceptions; a binary conversion is a separate setup.
Check the rubric, input length, answer options, and model configuration. A reranker adapted to pick an option and a native decision model can answer the same question through different serving paths.
Calibration asks whether reported probabilities match observed outcomes. Test an escalation threshold on your own labeled cases before letting a high-confidence answer approve an action or bypass a larger model.
Compare cost per completed decision alongside latency and accuracy. JevBench estimates some hosting costs from base-model reference prices and adjusts self-hosted latency. Those assumptions can change which system fits your budget.
The 85.71% is Perplexity’s reported accuracy across a fixed 7,210-row panel of 11 tasks from September 2026. Task samples have different sizes, and Decider was measured through the Perplexity API. The percentage measures correctness on those samples; it does not establish probability calibration or replace JevBench’s full composite.
Perplexity lists pplx-decider-v1-27b at $0.04 per million input tokens, with free output tokens and no per-request fee. State, images, and all questions count toward input usage. Use the returned input-token count to price a completed decision; a short label alone does not determine the request cost.
Decision models answer within a defined set of choices or score levels, often with probabilities your code can use. Reasoning models are evaluated on solving broader problems. A general-purpose model can act as a decision system, but its result depends on the prompt, adapter, and serving setup.
JevBench evaluates typed decision systems across intelligence, calibration, speed, and cost. DecisionBench from Hanno Labs publishes task-level accuracy and probability-quality measures, with applied decisions and reasoning in separate tracks. This page mirrors the pinned JevBench release and links to DecisionBench directly; the two scores are not interchangeable.
The category groups decision-system evidence for discovery. JevBench retains its official ranking and remains outside overall and category score calculations because its composite includes serving latency, pricing assumptions, and adapters. Model pages keep the evaluated configuration separate from its base model, so scores are not inherited between them.