# JevBench

> Typed-decision evaluation across 1,624 decisions, combining chance-corrected intelligence, probability calibration, latency, and cost.

Canonical page: https://benchlm.ai/benchmarks/jevbench

- Category: [Decision Models](/decision-models)
- Last updated: v1.5.4 · retrieved September 30, 2026

JevBench by [Florian Standhartinger](https://github.com/fstandhartinger) and contributors. [Original benchmark](https://github.com/fstandhartinger/jevbench) · [MIT license](/licenses/jevbench-mit.txt).

Copyright (c) 2026 Florian Standhartinger and contributors. The MIT notice applies to JevBench source code and method documentation. Aggregate results come from Benchmark Heaven's public versioned API. Evaluated model weights, external code, upstream datasets, and other third-party data retain their own terms. 

## About JevBench

- Year: 2026
- Tasks: 904 open decisions plus 720 sealed decisions per complete run
- Format: Official option A harmonic composite with published uncertainty
- Difficulty: Choice, Noul, and Score decisions with sealed generalization tests
- Paper: [JevBench v1.5 frozen method and disclosures](https://github.com/fstandhartinger/jevbench/blob/bb05a335bc809e61b20c0f745d25499a82b326fc/docs/METHOD-v1.5-README.md)

We mirror the complete v1.5.4 roster, including six unranked entries, and preserve every published option A score, rank, confidence interval, and adjacent-pair tie. The table keeps configurations separate and displays missing measurements explicitly. The original frozen method and its later headline amendment remain linked beside the data.

JevBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (112 systems, 106 ranked)

| Rank | System | Score | Intelligence | Calibration | Speed | Cost | Sealed Intelligence | USD / 1,000 decisions | Status | 95% composite interval | Published pairwise comparison |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | [Cygnet](/models/cygnet) | 73.70 | 71.09 | 87.01 | 90.97 | 56.43 | 69.72 | 0.02834 | ranked; estimate | 72.36–74.46 | Statistical tie with Winnow-12B Q8 |
| 2 | [Winnow-12B Q8](/models/winnow-12b) | 73.23 | 74.43 | 84.07 | 86.14 | 56.56 | 76.39 | 0.02806 | ranked; estimate | 72.02–73.99 | Published separation with Jev 1.13.0 |
| 3 | [Jev 1.13.0](/models/jev-1-13-0) | 72.13 | 72.00 | 88.03 | 83.81 | 54.73 | 72.45 | 0.03230 | ranked; API exposure; estimate | 71.01–72.61 | Not established |
| 4 | [JevK5 v0.3 4B](/models/jevk5-v0-3-4b) | 71.90 | 56.27 | 88.34 | 93.55 | 63.07 | 50.94 | 0.01702 | ranked; estimate | 69.39–72.95 | Not established |
| 5 | [Plumb-4B](/models/plumb-4b) | 71.56 | 55.85 | 87.44 | 93.45 | 63.07 | 50.81 | 0.01702 | ranked; estimate | 69.20–72.73 | Not established |
| 6 | [Jev-Omni](/models/jev-omni) | 71.50 | 70.49 | 82.60 | 84.68 | 56.05 | 70.50 | 0.02917 | ranked; estimate | 70.21–72.40 | Statistical tie with decider-4b v2 |
| 7 | [decider-4b v2](/models/decider-4b-v2) | 71.28 | 55.77 | 85.59 | 90.86 | 64.54 | 49.80 | 0.01520 | ranked; estimate | 69.11–72.35 | Not established |
| 8 | [Decision 4B v1.2](/models/decision-4b-v12) | 70.83 | 53.66 | 88.56 | 93.50 | 63.07 | 51.06 | 0.01702 | ranked; estimate | 68.52–71.97 | Not established |
| 9 | [Imajev-4B (RTX 5090)](/models/imajev-4b-rtx5090-a2) | 70.39 | 53.47 | 88.12 | 91.13 | 63.26 | 50.70 | 0.01677 | ranked; estimate | 67.80–71.61 | Not established |
| 10 | [Decision 4B v1.1](/models/decision-4b-v11) | 70.39 | 53.11 | 87.31 | 93.52 | 63.07 | 51.56 | 0.01702 | ranked; estimate | 66.86–71.57 | Not established |
| 11 | [Manchego v2.1](/models/manchego-v2-1) | 68.81 | 51.20 | 84.93 | 88.63 | 64.33 | 47.99 | 0.01545 | ranked; estimate | 59.64–70.28 | Not established |
| 12 | [SemIf Qwen3.5-4B](/models/semif-qwen3-5-4b) | 68.66 | 51.31 | 83.97 | 90.89 | 63.07 | 48.63 | 0.01702 | ranked; estimate | 60.06–70.16 | Statistical tie with spark-s1-4b-v6 |
| 13 | [spark-s1-4b-v6](/models/spark-s1-4b-v6) | 68.16 | 62.12 | 69.71 | 85.53 | 60.42 | 61.77 | 0.02086 | ranked; estimate | 66.27–69.70 | Statistical tie with metask-jev-4b |
| 14 | [metask-jev-4b](/models/metask-jev-4b) | 67.48 | 53.48 | 82.74 | 89.49 | 57.73 | 50.22 | 0.02564 | ranked; estimate | 65.39–68.69 | Statistical tie with Hopper |
| 15 | [Hopper](/models/hopper) | 67.47 | 49.86 | 87.87 | 87.23 | 62.26 | 51.02 | 0.01812 | ranked; estimate | 56.34–69.15 | Statistical tie with Malkuth-4B |
| 16 | [Malkuth-4B](/models/malkuth-4b) | 66.77 | 54.45 | 83.21 | 86.16 | 55.82 | 50.54 | 0.02968 | ranked; estimate | 64.94–67.89 | Not established |
| 17 | [Surogate Rune 26B-A4B v3 (RTX PRO 6000)](/models/surogate-rune-26b-a4b-v3-rtxpro6000-a2) | 66.47 | 69.74 | 88.30 | 85.97 | 48.97 | 70.53 | 0.05024 | ranked; estimate | 65.43–67.01 | Not established |
| 18 | [reflex 4B](/models/reflex-4b) | 65.24 | 51.49 | 86.84 | 68.77 | 63.14 | 47.78 | 0.01693 | ranked; estimate | 58.56–66.35 | Statistical tie with jev-local Qwen3.5-9B |
| 19 | [jev-local Qwen3.5-9B](/models/jev-local) | 65.24 | 56.28 | 77.82 | 72.98 | 58.86 | 54.37 | 0.02352 | ranked; estimate | 63.46–66.50 | Statistical tie with djev |
| 20 | [djev](/models/djev) | 64.16 | 72.33 | 80.41 | 91.00 | 48.22 | 71.95 | 0.05321 | ranked; estimate | 63.07–65.00 | Statistical tie with Raw Qwen3 4B Instruct 2507 direct logits |
| 21 | [Raw Qwen3 4B Instruct 2507 direct logits](/models/raw-qwen3-4b-instruct-2507) | 62.14 | 54.06 | 52.95 | 89.23 | 63.36 | 50.70 | 0.01665 | ranked; estimate | 59.35–64.22 | Statistical tie with jqv Qwen3-32B |
| 22 | [jqv Qwen3-32B](/models/jqv) | 60.77 | 49.09 | 86.77 | 83.36 | 51.17 | 46.83 | 0.04242 | ranked; estimate | 51.25–64.26 | Statistical tie with JevK5 v0.2.0 |
| 23 | [JevK5 v0.2.0](/models/jevk5-v0-2-0) | 58.11 | 46.70 | 84.87 | 90.86 | 63.07 | 43.40 | 0.01702 | ranked; estimate | 47.56–68.24 | Statistical tie with Qwen3.5-9B Jev-like data-mix v2 |
| 24 | [Qwen3.5-9B Jev-like data-mix v2](/models/qwen35-9b-jev-data-mix-v2) | 53.02 | 60.45 | 80.86 | 82.00 | 45.69 | 58.28 | 0.06461 | ranked; estimate | 51.88–53.88 | Published separation with Standard One 8B |
| 25 | [Standard One 8B](/models/standardone-8b) | 47.79 | 59.60 | 83.16 | 92.48 | 43.28 | 59.24 | 0.07772 | ranked; estimate | 46.67–48.54 | Statistical tie with NInfer Qwen3.8-Flash-Next mixed |
| 26 | [NInfer Qwen3.8-Flash-Next mixed](/models/ninfer-qwen3-8-flash-next) | 47.47 | 67.20 | 88.54 | 88.61 | 42.53 | 64.14 | 0.08233 | ranked; estimate | 46.70–47.88 | Not established |
| 27 | [Instinct Dual 4B](/models/instinct-dual-4b) | 46.97 | 43.12 | 88.30 | 81.95 | 60.19 | 39.95 | 0.02124 | ranked; API exposure; estimate | 37.86–56.77 | Not established |
| 28 | [swanOne](/models/swanone) | 46.56 | 71.22 | 87.06 | 84.74 | 42.15 | 71.95 | 0.08478 | ranked; estimate | 45.81–46.94 | Statistical tie with Raw Qwen3 8B direct logits |
| 29 | [Raw Qwen3 8B direct logits](/models/raw-qwen3-8b) | 45.22 | 51.14 | 49.19 | 86.84 | 45.54 | 45.50 | 0.06539 | ranked; estimate | 36.89–46.74 | Statistical tie with decider-2b |
| 30 | [decider-2b](/models/decider-2b) | 45.10 | 42.34 | 71.53 | 94.40 | 64.92 | 35.89 | 0.01477 | ranked; estimate | 33.82–54.26 | Statistical tie with system-one Qwen3-8B |
| 31 | [system-one Qwen3-8B](/models/system-one-sg) | 44.14 | 50.47 | 49.38 | 90.88 | 44.97 | 47.41 | 0.06829 | ranked; estimate | 36.79–45.66 | Statistical tie with system-one-open |
| 32 | [system-one-open](/models/system-one-open) | 42.44 | 41.64 | 71.63 | 78.27 | 68.33 | 38.85 | 0.01137 | ranked; API exposure; estimate | 33.84–51.92 | Statistical tie with Autoloops Gemma 4 31B IT |
| 33 | [Autoloops Gemma 4 31B IT](/models/autoloops-gemma-4-31b-it) | 40.53 | 76.68 | 85.84 | 83.92 | 39.59 | 76.49 | 0.10323 | ranked; API exposure; tariff | 40.09–40.76 | Statistical tie with GPT-6 Luna (low) |
| 34 | [GPT-6 Luna (low)](/models/gpt-6-luna-low) | 40.48 | 95.29 | 94.92 | 73.20 | 39.06 | 95.86 | 0.10750 | ranked; API exposure; estimate | 40.27–40.66 | Published separation with GPT-6 Luna (medium) |
| 35 | [GPT-6 Luna (medium)](/models/gpt-6-luna-medium) | 38.75 | 96.24 | 95.56 | 73.19 | 38.32 | 96.15 | 0.11379 | ranked; API exposure; estimate | 38.57–38.93 | Statistical tie with JevOne |
| 36 | [JevOne](/models/jevone) | 38.25 | 54.13 | 84.94 | 90.43 | 39.84 | 54.36 | 0.10122 | ranked; estimate | 37.38–38.79 | Statistical tie with kev 4B |
| 37 | [kev 4B](/models/kev-4b) | 38.07 | 39.86 | 67.64 | 85.48 | 65.78 | 33.20 | 0.01382 | ranked; estimate | 27.48–46.77 | Statistical tie with kev 8B |
| 38 | [kev 8B](/models/kev-8b) | 34.15 | 48.30 | 71.02 | 84.14 | 40.42 | 42.28 | 0.09682 | ranked; estimate | 26.79–37.48 | Statistical tie with open-alternative-jev Qwen3.5-4B |
| 39 | [open-alternative-jev Qwen3.5-4B](/models/open-alternative-jev) | 33.56 | 37.36 | 76.87 | 91.29 | 63.28 | 32.12 | 0.01675 | ranked; estimate | 25.75–41.66 | Statistical tie with Bespoke Nimble 9B |
| 40 | [Bespoke Nimble 9B](/models/nimble-9b) | 31.82 | 63.75 | 77.16 | 82.95 | 36.75 | 62.35 | 0.12833 | ranked; estimate | 31.19–32.27 | Statistical tie with Malkuth-2B |
| 41 | [Malkuth-2B](/models/malkuth-2b) | 29.90 | 35.54 | 75.19 | 91.80 | 65.57 | 28.61 | 0.01405 | ranked; estimate | 20.96–38.98 | Statistical tie with openjev-sglang Qwen3.6-35B-A3B |
| 42 | [openjev-sglang Qwen3.6-35B-A3B](/models/openjev-sglang) | 29.03 | 58.65 | 82.87 | 78.07 | 35.63 | 55.94 | 0.13981 | ranked; API exposure; estimate | 28.42–29.40 | Published separation with decider-35b-a3b |
| 43 | [decider-35b-a3b](/models/decider-35b-a3b) | 27.50 | 60.46 | 82.04 | 91.00 | 34.39 | 56.78 | 0.15387 | ranked; estimate | 26.93–27.87 | Statistical tie with local-jev Qwen3.5-4B |
| 44 | [local-jev Qwen3.5-4B](/models/localjev-qwen3-5-4b) | 25.82 | 33.75 | 82.45 | 83.74 | 59.27 | 30.79 | 0.02279 | ranked; estimate | 19.89–32.96 | Not established |
| 45 | [Nemotron Diffusion 8B (optimized vLLM)](/models/nemotron-diffusion-8b) | 25.68 | 33.84 | 76.25 | 93.74 | 55.49 | 28.15 | 0.03045 | ranked; estimate | 18.60–32.78 | Not established |
| 46 | [Open-Jev 9B](/models/open-jev-zefan-9b) | 24.36 | 63.82 | 81.53 | 73.59 | 33.05 | 65.86 | 0.17041 | ranked; estimate | 23.93–24.64 | Statistical tie with Decision 2B v59 |
| 47 | [Decision 2B v59](/models/decision-2b) | 22.52 | 31.31 | 86.22 | 90.36 | 66.37 | 27.98 | 0.01321 | ranked; estimate | 16.68–28.77 | Statistical tie with GPT-5.6 Luna (low) |
| 48 | [GPT-5.6 Luna (low)](/models/gpt-5-6-luna-low) | 22.37 | 94.33 | 94.71 | 74.06 | 30.67 | 95.51 | 0.20467 | ranked; API exposure; tariff | 22.26–22.46 | Statistical tie with typecastlm |
| 49 | [typecastlm](/models/typecastlm) | 21.83 | 31.24 | 76.71 | 92.04 | 63.96 | 38.78 | 0.01590 | ranked; estimate | 16.23–28.58 | Statistical tie with JEV Qwen3.5-9B Base NVFP4 |
| 50 | [JEV Qwen3.5-9B Base NVFP4](/models/jev-qwen3-5-9b-base-nvfp4) | 20.07 | 32.24 | 80.93 | 93.54 | 47.59 | 31.61 | 0.05584 | ranked; estimate | 15.19–25.51 | Statistical tie with Gemini 3.1 Flash-Lite |
| 51 | [Gemini 3.1 Flash-Lite](/models/gemini-3-1-flash-lite) | 19.58 | 77.59 | 74.68 | 80.05 | 29.76 | 78.56 | 0.21941 | ranked; API exposure; tariff | 19.30–19.83 | Not established |
| 52 | [AutoJev-27B](/models/autojev-27b) | 19.54 | 72.78 | 87.70 | 87.62 | 29.37 | 73.65 | 0.22619 | ranked; estimate | 19.30–19.63 | Not established |
| 53 | [AutoJev-27B (RTX PRO 6000)](/models/autojev-27b-rtxpro6000-a2) | 19.48 | 72.75 | 86.68 | 87.07 | 29.37 | 73.49 | 0.22619 | ranked; estimate | 19.26–19.60 | Not established |
| 54 | [NInfer Qwen3.8-27B NVFP4](/models/ninfer-qwen3-8-27b) | 18.74 | 65.47 | 85.91 | 89.86 | 29.12 | 63.97 | 0.23052 | ranked; estimate | 18.45–18.92 | Not established |
| 55 | [Eikos-27B](/models/eikos-27b) | 18.50 | 75.08 | 86.29 | 87.63 | 28.69 | 76.02 | 0.23824 | ranked; estimate | 18.28–18.60 | Not established |
| 56 | [NInfer Qwen3.8-27B NVFP4 (T=1.5)](/models/ninfer-qwen3-8-27b-t1-5) | 18.48 | 61.08 | 86.49 | 89.86 | 29.12 | 59.72 | 0.23052 | ranked; estimate | 18.18–18.65 | Statistical tie with Instinct Qwen3.8-27B |
| 57 | [Instinct Qwen3.8-27B](/models/instinct) | 18.32 | 62.75 | 84.96 | 81.68 | 29.16 | 63.58 | 0.22979 | ranked; API exposure; estimate | 18.03–18.51 | Published separation with OpenJev (thinking, BF16) |
| 58 | [OpenJev (thinking, BF16)](/models/openjev-thinking) | 17.94 | 84.21 | 83.10 | 73.86 | 28.51 | 85.90 | 0.24146 | ranked; estimate | 17.76–18.08 | Published separation with djev (thinking) |
| 59 | [djev (thinking)](/models/djev-thinking) | 17.40 | 77.35 | 95.66 | 72.32 | 28.13 | 78.14 | 0.24866 | ranked; estimate | 17.28–17.51 | Published separation with LitJev Qwen3.8-27B |
| 60 | [LitJev Qwen3.8-27B](/models/litjev) | 16.31 | 58.32 | 84.45 | 68.22 | 28.36 | 58.59 | 0.24438 | ranked; estimate | 16.03–16.48 | Not established |
| 61 | [Bev / Bonsai 27B](/models/bev-bonsai-27b) | 15.77 | 53.08 | 77.68 | 72.73 | 28.23 | 57.46 | 0.24676 | ranked; estimate | 15.13–16.05 | Not established |
| 62 | [Raw Phi-4 mini direct logits](/models/raw-phi-4-mini) | 15.22 | 27.62 | 71.41 | 89.17 | 53.22 | 19.96 | 0.03625 | ranked; estimate | 10.26–20.74 | Statistical tie with OpenSourceJev Qwen3.5-4B Q4_K_M |
| 63 | [OpenSourceJev Qwen3.5-4B Q4_K_M](/models/opensourcejev-qwen35-4b-q4km) | 13.25 | 25.77 | 76.04 | 72.51 | 69.14 | 23.68 | 0.01069 | ranked; estimate | 8.98–18.31 | Statistical tie with reflex-27b |
| 64 | [reflex-27b](/models/reflex-27b) | 13.19 | 62.79 | 85.93 | 69.20 | 25.80 | 62.50 | 0.29734 | ranked; estimate | 12.99–13.30 | Published separation with Open-Jev 2B |
| 65 | [Open-Jev 2B](/models/open-jev-zefan-2b) | 9.07 | 33.56 | 73.62 | 75.76 | 33.05 | 28.34 | 0.17041 | ranked; estimate | 6.62–11.54 | Statistical tie with GLiNER2 large |
| 66 | [GLiNER2 large](/models/gliner2-large) | 8.40 | 22.49 | 42.38 | 64.78 | 77.59 | 18.96 | 0.00558 | ranked; estimate | 5.21–12.34 | Statistical tie with Qwen3-Reranker-4B |
| 67 | [Qwen3-Reranker-4B](/models/qwen3-reranker-4b) | 7.06 | 21.03 | 76.26 | 79.61 | 48.40 | 13.36 | 0.05249 | ranked; estimate | 4.26–10.74 | Statistical tie with DeepSeek V4.1 Flash |
| 68 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | 6.65 | 93.69 | 96.92 | 69.40 | 19.09 | 95.53 | 0.49760 | ranked; API exposure; estimate | 6.62–6.66 | Statistical tie with SimpleJev Qwen3.5-0.8B |
| 69 | [SimpleJev Qwen3.5-0.8B](/models/simplejev-qwen3-5-0-8b) | 4.00 | 16.75 | 46.67 | 59.07 | 70.31 | 20.16 | 0.00976 | ranked; estimate | 2.06–6.66 | Statistical tie with SimpleJev Qwen3.8-27B |
| 70 | [SimpleJev Qwen3.8-27B](/models/simplejev-qwen3-8-27b) | 3.36 | 72.75 | 87.20 | 74.92 | 14.90 | 73.60 | 0.68680 | ranked; API exposure; estimate | 3.33–3.37 | Statistical tie with decision-machine-1 |
| 71 | [decision-machine-1](/models/decision-machine-1) | 3.24 | 14.82 | 81.18 | 92.77 | 56.30 | 5.13 | 0.02863 | ranked; API exposure; estimate | 1.73–5.39 | Statistical tie with GLiNER2.5 multi |
| 72 | [GLiNER2.5 multi](/models/gliner2-5-multi) | 2.72 | 14.00 | 57.96 | 66.57 | 86.62 | 12.21 | 0.00279 | ranked; estimate | 1.20–5.10 | Not established |
| 73 | [Bosun v3.1 0.6B](/models/bosun-v31-0-6b) | 2.50 | 13.58 | 65.16 | 61.76 | 77.55 | 5.33 | 0.00560 | ranked; estimate | 1.15–4.56 | Not established |
| 74 | [GLiNER2.5 base](/models/gliner2-5-base) | 2.27 | 13.48 | 35.57 | 70.33 | 86.62 | 9.71 | 0.00279 | ranked; estimate | 0.90–4.50 | Not established |
| 75 | [Deem 0.8B v1](/models/deem-0-8-v1) | 2.14 | 13.02 | 38.88 | 84.32 | 80.29 | 3.59 | 0.00454 | ranked; estimate | 0.90–4.34 | Not established |
| 76 | [JevAct](/models/jevact) | 1.52 | 11.24 | 62.54 | 76.43 | 68.56 | 3.47 | 0.01117 | ranked; API exposure; estimate | 0.52–3.18 | Statistical tie with CLM-8B |
| 77 | [CLM-8B](/models/clm-8b) | 1.51 | 11.43 | 48.48 | 93.29 | 50.50 | 5.77 | 0.04466 | ranked; estimate | 0.48–3.17 | Statistical tie with kev 0.6B |
| 78 | [kev 0.6B](/models/kev-0-6b) | 1.26 | 10.34 | 67.65 | 87.22 | 80.09 | -1.49 | 0.00461 | ranked; estimate | 0.49–2.50 | Statistical tie with Raw Qwen3 0.6B direct logits |
| 79 | [Raw Qwen3 0.6B direct logits](/models/raw-qwen3-0-6b) | 1.08 | 10.56 | 21.45 | 90.39 | 77.58 | 11.11 | 0.00559 | ranked; estimate | 0.27–2.61 | Statistical tie with GLiNER2.5 small |
| 80 | [GLiNER2.5 small](/models/gliner2-5-small) | 0.89 | 9.19 | 55.90 | 77.36 | 86.62 | 1.65 | 0.00279 | ranked; estimate | 0.23–2.20 | Statistical tie with Raw Qwen3 1.7B direct logits |
| 81 | [Raw Qwen3 1.7B direct logits](/models/raw-qwen3-1-7b) | 0.85 | 9.67 | 21.57 | 90.22 | 68.55 | 7.08 | 0.01118 | ranked; estimate | 0.15–2.28 | Statistical tie with Mirror |
| 82 | [Mirror](/models/mirror) | 0.22 | 5.60 | 43.21 | 63.64 | 89.27 | 4.45 | 0.00228 | ranked; estimate | 0.01–0.90 | Statistical tie with ZeroEntropy zerank-2 |
| 83 | [ZeroEntropy zerank-2](/models/zerank-2) | 0.13 | 4.76 | 81.71 | 80.28 | 48.40 | -0.44 | 0.05249 | ranked; estimate | 0.01–0.49 | Statistical tie with jeff |
| 84 | [jeff](/models/jeff) | 0.12 | 4.43 | 80.27 | 55.76 | 81.09 | -3.56 | 0.00427 | ranked; estimate | 0.00–0.51 | Statistical tie with smalljev semantic-v9 |
| 85 | [smalljev semantic-v9](/models/smalljev) | 0.07 | 3.66 | 73.09 | 86.46 | 60.77 | -3.52 | 0.02030 | ranked; estimate | 0.00–0.38 | Not established |
| 86 | [Laya multilingual](/models/laya-multilingual) | 0.02 | 2.35 | 43.58 | 73.72 | 82.16 | 1.92 | 0.00393 | ranked; estimate | 0.00–0.27 | Not established |
| 87 | [OpenDecision ModernBERT-large](/models/opendecision) | 0.01 | 2.07 | 72.52 | 86.58 | 79.11 | -5.26 | 0.00497 | ranked; estimate | 0.00–0.19 | Statistical tie with BAAI bge-reranker-v2-m3 |
| 88 | [BAAI bge-reranker-v2-m3](/models/bge-reranker-v2-m3) | 0.00 | 0.00 | 83.33 | 90.55 | 59.33 | -22.80 | 0.02269 | ranked; estimate | 0.00–0.00 | Statistical tie with Certo v1 |
| 89 | [Certo v1](/models/certo) | 0.00 | 0.00 | 87.52 | 91.43 | 96.91 | -24.93 | 0.00127 | ranked; estimate | 0.00–0.00 | Statistical tie with Decision Fast v53a |
| 90 | [Decision Fast v53a](/models/decision-fast) | 0.00 | 0.00 | 76.08 | 91.44 | 80.09 | -11.70 | 0.00461 | ranked; estimate | 0.00–0.00 | Statistical tie with Alibaba GTE Reranker ModernBERT-base |
| 91 | [Alibaba GTE Reranker ModernBERT-base](/models/gte-reranker-modernbert-base) | 0.00 | 0.00 | 75.20 | 91.49 | 69.27 | -23.16 | 0.01058 | ranked; estimate | 0.00–0.00 | Statistical tie with kev 0.5B |
| 92 | [kev 0.5B](/models/kev-0-5b) | 0.00 | 0.00 | 64.98 | 88.45 | 80.09 | -22.26 | 0.00461 | ranked; estimate | 0.00–0.00 | Statistical tie with Laya |
| 93 | [Laya](/models/laya) | 0.00 | 0.00 | 73.71 | 73.89 | 84.89 | -13.19 | 0.00319 | ranked; estimate | 0.00–0.00 | Statistical tie with lev-350m |
| 94 | [lev-350m](/models/lev-350m) | 0.00 | 0.00 | 77.68 | 93.63 | 80.04 | -18.01 | 0.00463 | ranked; estimate | 0.00–0.00 | Statistical tie with Qwen3.5-0.8B Decision Model |
| 95 | [Qwen3.5-0.8B Decision Model](/models/mghafiri-qwen3-5-0-8b-decision-model) | 0.00 | 0.00 | 73.57 | 71.82 | 79.52 | -3.71 | 0.00482 | ranked; estimate | 0.00–0.02 | Statistical tie with Mixedbread mxbai-rerank-base-v2 |
| 96 | [Mixedbread mxbai-rerank-base-v2](/models/mxbai-rerank-base-v2) | 0.00 | 0.00 | 86.61 | 89.34 | 60.34 | -24.04 | 0.02100 | ranked; estimate | 0.00–0.00 | Statistical tie with Needle 3 (2-bit) |
| 97 | [Needle 3 (2-bit)](/models/needle-3) | 0.00 | 0.00 | 0.00 | 34.13 | 61.57 | -21.11 | 0.01910 | ranked; estimate | 0.00–0.00 | Statistical tie with Needle 3 (options as tools) |
| 98 | [Needle 3 (options as tools)](/models/needle-3-tools) | 0.00 | 0.00 | 0.00 | 41.00 | 61.57 | -19.90 | 0.01910 | ranked; estimate | 0.00–0.00 | Statistical tie with open-jev-deberta-v3-large |
| 99 | [open-jev-deberta-v3-large](/models/open-jev-deberta-v3-large) | 0.00 | 0.00 | 77.06 | 68.25 | 77.59 | -14.61 | 0.00558 | ranked; estimate | 0.00–0.00 | Statistical tie with Open Jev JSON Canvas |
| 100 | [Open Jev JSON Canvas](/models/open-jev-json-canvas-joshuasp) | 0.00 | 77.14 | 0.00 | 85.57 | 49.30 | 75.12 | 0.04900 | ranked; estimate | 0.00–0.00 | Statistical tie with openJev Verdict |
| 101 | [openJev Verdict](/models/openjev-verdict) | 0.00 | 0.00 | 52.22 | 83.88 | 86.62 | -14.09 | 0.00279 | ranked; estimate | 0.00–0.01 | Statistical tie with openJev Verdict 1.4 |
| 102 | [openJev Verdict 1.4](/models/openjev-verdict-1-4) | 0.00 | 0.00 | 80.33 | 80.64 | 86.62 | -18.28 | 0.00279 | ranked; estimate | 0.00–0.00 | Statistical tie with Qwen3.8-27B (Chutes TEE) |
| 103 | [Qwen3.8-27B (Chutes TEE)](/models/qwen3-8-27b-chutes) | 0.00 | 95.56 | 98.09 | 56.85 | 0.00 | 98.23 | 2.17839 | ranked; API exposure; estimate | 0.00–0.00 | Statistical tie with verdict-small |
| 104 | [verdict-small](/models/verdict-small) | 0.00 | 0.00 | 59.45 | 82.06 | 100.00 | -2.95 | 0.00087 | ranked; estimate | 0.00–0.01 | Statistical tie with Von 395M |
| 105 | [Von 395M](/models/von-395m) | 0.00 | 0.00 | 83.48 | 75.72 | 82.71 | -16.55 | 0.00377 | ranked; estimate | 0.00–0.00 | Not established |
| 106 | [Laya typed-decisions](/models/laya-typed-decisions) | 0.00 | 0.00 | 83.34 | 63.38 | 82.53 | -16.92 | 0.00382 | ranked; estimate | 0.00–0.00 | Not established |
| Unranked | [classifier.dev (fast)](/models/classifier-dev-fast) | 74.66 | 75.81 | 89.19 | 81.99 | 58.89 | 77.14 | 0.02345 | honorable_mention; API exposure; estimate; runs on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2) | 73.50–75.20 | Not established |
| Unranked | [SimpleJev Qwen3.6-35B-A3B](/models/simplejev-qwen3-6-35b-a3b) | 0.00 | 0.00 | 78.50 | 75.18 | 35.15 | -22.97 | 0.14512 | partial; API exposure; estimate; Partial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked. | 0.00–0.00 | Not established |
| Unranked | [Decision-4B (Eval Engine / Chromia)](/models/evalengine-decision-4b) | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | 0.01372 | unranked; estimate; PARTIAL / UNRANKED: 1,550 of 1,624 valid answers; 74 context overflows in our evaluator at 2,048 tokens. No official score or rank. | Not measured | Not established |
| Unranked | [Jobe Qwen3.5-4B](/models/jobe-qwen3-5-4b) | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | not measured; Not measured in JevBench v1.5; no current score or rank. | Not measured | Not established |
| Unranked | [Mica v0.1 4B](/models/mica-v0-1-4b) | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | partial run; The frozen refusal policy stopped the run after 1,088 of 1,624 rows: 27 documented refusals were mapped to HTTP 422, then three consecutive passthrough HTTP 400 refusals triggered exit 6. The remaining 536 rows have no scores, so this system is not eligible for an official rank. | Not measured | Not established |
| Unranked | [OpenJev DiffusionGemma 26B-A4B NVFP4](/models/openjev-razorback16) | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | not measured; Not measured in JevBench v1.5; no current score or rank. | Not measured | Not established |

Option A is the official headline. Composite intervals and pairwise ties follow the published artifact; absent markers establish neither a tie nor a separation. Sealed Intelligence is chance-corrected competence, not percent accuracy.

Cost is USD per 1,000 whole decisions. Estimated costs use the published benchmark reference; they are not API token tariffs. API exposure means the operator received sealed item text without answers.

## How JevBench changed

This page follows the latest imported release, currently v1.5.4. Earlier results remain labeled by version on model pages. Changes to the tasks and scoring prevent a direct comparison of scores across the 1.4 and 1.5 protocols.

| Release | Decisions per complete run | What changed |
|---|---|---|
| [v1.5.4 · current imported release](https://benchmarkheaven.com/api/jevbench/v1.5.4) | 904 open + 720 sealed | The roster has 112 systems and 106 official ranks. Manchego v2.1 joins after a complete run; earlier scores, intervals, and the scoring method remain unchanged. |
| [v1.5 · protocol change](https://github.com/fstandhartinger/jevbench/blob/bb05a335bc809e61b20c0f745d25499a82b326fc/docs/METHOD-v1.5-ADDENDUM-HEADLINE-A-EQUAL-TYPES.md) | 904 open + 720 sealed | Choice, Noul, and Score receive equal weight, and sealed results carry half of Intelligence. The official headline changed to option A after the method owner reviewed results. Published intervals and paired statistical ties help distinguish small gaps. |
| [v1.4.2.2 · September 27, 2026](https://github.com/fstandhartinger/jevbench/blob/5b405e7e4efff2cfbbb2fe67db9b0cbad0704b8d/results/v1.4.2.2/jevbench-v1.4.2.2-results.json) | 534 frozen + 308 sealed | Imajev-4B joined a roster of 95 systems, with 91 ranked. Version 1.4 used a harmonic composite and a 20% sealed share of Intelligence; later roster additions preserved earlier measurements. |

## FAQ

### What does JevBench 1.5 measure?

JevBench 1.5 evaluates typed decisions using Choice, Noul, and Score requests across 904 open and 720 sealed items. Its current official option A combines Intelligence, Calibration, Speed, and Cost equally. The published table includes composite confidence intervals and statistical ties; its score is an index rather than percent accuracy.

### Why are some JevBench 1.5 systems unranked?

The v1.5.4 roster contains 112 systems and 106 official ranks. classifier.dev remains an honorable mention because it runs another entrant. Two entries have partial evaluations, and three are incomplete or unmeasured. We retain their published evidence and reasons while leaving unavailable scores and ranks missing.

### Are JevBench 1.4 and 1.5 scores comparable?

The versions use different task sets and scoring. Version 1.5 doubles the sealed share of Intelligence to 50%, adds native evaluation of three decision types, and publishes uncertainty. Model evidence rows retain their versioned benchmark keys, and this page explains the earlier releases. A higher score on 1.5 does not by itself establish that a model improved.
