# Decision Models

Compare AI systems that choose an allowed answer and report probabilities for routing, classification, and scoring. JevBench keeps decision quality, confidence, latency, and cost visible so you can check the trade-offs for your application.

Canonical page: https://benchlm.ai/decision-models

JevBench v1.5.4 · retrieved September 30, 2026: 112 systems, 106 ranked.

## Typed decisions your code can use

Decision models take application state, a written question, and an allowed answer space. Their output fits a check your code can act on. A bounded answer can still be wrong, and reported confidence needs testing on the decisions your application actually makes.

| Primitive | Example task | Output |
| --- | --- | --- |
| Choice | Route a support ticket to billing, technical support, or account support. | An option and its probability distribution. |
| Noul | Check whether a request meets a written approval rule. | A probability that a yes/no proposition holds. |
| Score | Rate a bug against an ordered severity rubric. | A score and probabilities over the allowed levels. |

[OpenRouter decision-model explainer](https://openrouter.ai/blog/insights/what-is-jev/)

## Perplexity Decider v1 27B

Perplexity released Decider v1 27B on October 1, 2026, as a Qwen3.8-27B fine-tune with Apache 2.0 weights. It reads text, JSON, and images, then returns yes/no probabilities, choices, or rubric scores. The hosted Decisions API charges $0.04 per million input tokens, with free output and no per-request fee.

[Model profile and pricing](/models/pplx-decider-v1-27b) · [Perplexity Decisions API](https://docs.perplexity.ai/docs/decisions/quickstart) · [Official price](https://docs.perplexity.ai/docs/getting-started/pricing#decisions-api-pricing) · [Apache 2.0 weights](https://huggingface.co/perplexity-ai/pplx-decider-v1-27b)

## GLiDE on Decision Index 0.2.1

Fastino reports 64.81 skill points for GLiDE on Decision Index 0.2.1, compared with 57.91 for the published Jev 1.13.0 reference. Its September 30, 2026 announcement uses the official scorer; GLiDE is not listed on the public board.

[GLiDE profile and pricing](/models/fastino-glide) · [Fastino API reference](https://docs.fastino.ai/inference/systemone) · [Official GLiDE price](https://docs.fastino.ai/pricing)

| Metric | Unit | GLiDE | Jev 1.13.0 reference |
| --- | --- | ---: | ---: |
| [Decision Index 0.2.1](/benchmarks/decision-index-0-2-1-fastino) | skill points | 64.81 | 57.91 |
| [Knowledge and Reasoning](/benchmarks/decision-index-knowledge-fastino) | skill points | 62.9 | 51.4 |
| [Language Understanding](/benchmarks/decision-index-language-fastino) | skill points | 63.8 | 62.0 |
| [Retrieval and Classification](/benchmarks/decision-index-retrieval-fastino) | skill points | 60.9 | 55.4 |
| [Tools and Automation](/benchmarks/decision-index-tools-fastino) | skill points | 83.5 | 75.1 |
| [Arts and Human Taste](/benchmarks/decision-index-arts-fastino) | skill points | 46.0 | 37.7 |
| [CLadder](/benchmarks/cladder-decision-index-fastino) | % | 88.7% | 72.6% |
| [CRUXEval](/benchmarks/cruxeval-decision-index-fastino) | % | 92.6% | 73.0% |

### Other overall references in Fastino’s launch chart

| Evaluated system | Skill points | Source qualification |
| --- | ---: | --- |
| [d1](/models/liquid-d1) | 58.9 | Liquid AI self-reported reproduction, quoted by Fastino |
| [Surogate Rune 26B-A4B v3](/models/decision-index-rune-26b-a4b-v3) | 57.44 | Published Decision Index reference quoted by Fastino |
| [Decider chat · Gemma-4-31B](/models/decision-index-decider-chat-gemma4-31b) | 57.33 | Inference technique on Gemma-4-31B; published reference quoted by Fastino |
| [AutoJev-27B](/models/decision-index-autojev-27b) | 56.40 | Published Decision Index reference quoted by Fastino |
| [simple-jev · Qwen3.8-27B](/models/decision-index-simple-jev-qwen3-8-27b) | 55.74 | Featherless inference technique; published reference quoted by Fastino |

The five area scores and overall index are chance-adjusted skill points. CLadder and CRUXEval are raw accuracies. Fastino reports a complete run, but publishes exact GLiDE values for only these eight metrics. Radar-chart gaps do not supply the other individual benchmark accuracies. We have not rerun the evaluation, and these results stay outside general model rankings. The launch chart also quotes five other overall references, including Liquid AI’s self-reported d1 reproduction; their category and task scores are not inferred.

[Fastino launch and results](https://fastino.ai/blog/introducing-glide-the-first-thinking-decision-model)

## Perplexity’s 11-task decision panel

Perplexity reports 85.71% accuracy for Decider v1 27B, 84.51% for TypeSafe Jev 1.13.0, and 74.76% for Qwen3.8-27B on a fixed 7,210-row panel. The chart covers September 2026 and was published October 1. Decider was measured through the Perplexity API.

These are provider-reported results. The task samples have different sizes, so the published overall weights them differently. Prompts, sampled item identities, label mappings, baseline inference settings, and uncertainty are not specified in the model card. This panel stays separate from full-dataset results and general model rankings; JevBench public-hard accuracy is separate from its full composite.

| Benchmark | Rows | [Jev 1.13.0](/models/jev-1-13-0) | [Qwen3.8-27B](/models/qwen3-8-27b) | [Perplexity Decider](/models/pplx-decider-v1-27b) |
| --- | ---: | ---: | ---: | ---: |
| [WinoGrande](/benchmarks/winogrande-perplexity-panel) | 1,000 | 90.70% | 73.10% | 83.30% |
| [FinancialPhraseBank](/benchmarks/financialphrasebank-perplexity-panel) | 999 | 76.98% | 75.68% | 84.18% |
| [RAGTruth](/benchmarks/ragtruth-perplexity-panel) | 1,500 | 77.27% | 61.53% | 88.80% |
| [JudgeBench](/benchmarks/judgebench-perplexity-panel) | 350 | 78.57% | 68.86% | 78.29% |
| [BBH](/benchmarks/bbh-perplexity-panel) | 750 | 94.27% | 72.80% | 82.80% |
| [JevBench public hard](/benchmarks/jevbench-public-hard-perplexity-panel) | 101 | 73.27% | 72.28% | 70.30% |
| [TabFact](/benchmarks/tabfact-perplexity-panel) | 500 | 89.80% | 78.60% | 90.60% |
| [ContractNLI](/benchmarks/contractnli-perplexity-panel) | 510 | 77.45% | 80.78% | 80.78% |
| [Circa](/benchmarks/circa-perplexity-panel) | 500 | 84.60% | 87.00% | 89.20% |
| [Belebele](/benchmarks/belebele-perplexity-panel) | 500 | 95.00% | 93.20% | 94.00% |
| [TruthfulQA binary](/benchmarks/truthfulqa-binary-perplexity-panel) | 500 | 92.00% | 82.80% | 85.40% |
| [Overall](/benchmarks/perplexity-decision-panel) | 7,210 | 84.51% | 74.76% | 85.71% |

[Panel setup and limits](/benchmarks/perplexity-decision-panel) · [Pinned Perplexity model card](https://huggingface.co/perplexity-ai/pplx-decider-v1-27b/blob/9ce1abcf1f00209405376b5bcc81225c8f8cf514/README.md) · [Official launch chart](https://x.com/perplexitydevs/status/2105725598882832414)

## JevBench decision-model leaderboard

1624 decisions. The official composite is an index rather than percent accuracy and stays outside general rankings. Option A is the reviewed headline, selected after the method owner saw the results. Costs retain source tariffs or estimates. Individual intervals and adjacent-pair ties remain separate claims.

| Rank | System | Score | Intelligence | Calibration | Speed | Cost | 95% score interval | Adjacent pair |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | [Cygnet (blockbrain, frozen Gemma-4-12B-it)](/models/cygnet) | 73.70 | 71.09 | 87.01 | 90.97 | 56.43 | 72.36–74.46 | Statistical tie with Winnow-12B Q8 |
| 2 | [Winnow-12B Q8](/models/winnow-12b) | 73.23 | 74.43 | 84.07 | 86.14 | 56.56 | 72.02–73.99 | Published separation with Jev 1.13.0 |
| 3 | [Jev 1.13.0 (TypeSafe AI)](/models/jev-1-13-0) | 72.13 | 72.00 | 88.03 | 83.81 | 54.73 | 71.01–72.61 | Not established |
| 4 | [JevK5 v0.3 (4B)](/models/jevk5-v0-3-4b) | 71.90 | 56.27 | 88.34 | 93.55 | 63.07 | 69.39–72.95 | Not established |
| 5 | [Plumb-4B (crh225, JevK5 v0.2 + LoRA)](/models/plumb-4b) | 71.56 | 55.85 | 87.44 | 93.45 | 63.07 | 69.20–72.73 | Not established |
| 6 | [Jev-Omni (akhilaaa3, Gemma-4-12B merged)](/models/jev-omni) | 71.50 | 70.49 | 82.60 | 84.68 | 56.05 | 70.21–72.40 | Statistical tie with decider-4b v2 |
| 7 | [decider-4b v2 (Mapika)](/models/decider-4b-v2) | 71.28 | 55.77 | 85.59 | 90.86 | 64.54 | 69.11–72.35 | Not established |
| 8 | [Decision 4B v1.2 (FlyMyJev, Qwen3.5-4B + LoRA)](/models/decision-4b-v12) | 70.83 | 53.66 | 88.56 | 93.50 | 63.07 | 68.52–71.97 | Not established |
| 9 | [Imajev-4B (RTX 5090)](/models/imajev-4b-rtx5090-a2) | 70.39 | 53.47 | 88.12 | 91.13 | 63.26 | 67.80–71.61 | Not established |
| 10 | [Decision 4B v1.1 (FlyMyJev, Qwen3.5-4B + LoRA)](/models/decision-4b-v11) | 70.39 | 53.11 | 87.31 | 93.52 | 63.07 | 66.86–71.57 | Not established |
| 11 | [Manchego v2.1](/models/manchego-v2-1) | 68.81 | 51.20 | 84.93 | 88.63 | 64.33 | 59.64–70.28 | Not established |
| 12 | [SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)](/models/semif-qwen3-5-4b) | 68.66 | 51.31 | 83.97 | 90.89 | 63.07 | 60.06–70.16 | Statistical tie with spark-s1-4b-v6 |
| 13 | [spark-s1-4b-v6 (Open Spark Jev, abhishek085)](/models/spark-s1-4b-v6) | 68.16 | 62.12 | 69.71 | 85.53 | 60.42 | 66.27–69.70 | Statistical tie with metask-jev-4b |
| 14 | [metask-jev-4b](/models/metask-jev-4b) | 67.48 | 53.48 | 82.74 | 89.49 | 57.73 | 65.39–68.69 | Statistical tie with Hopper |
| 15 | [Hopper](/models/hopper) | 67.47 | 49.86 | 87.87 | 87.23 | 62.26 | 56.34–69.15 | Statistical tie with Malkuth-4B |
| 16 | [Malkuth-4B (newfull5, Kev post-train)](/models/malkuth-4b) | 66.77 | 54.45 | 83.21 | 86.16 | 55.82 | 64.94–67.89 | Not established |
| 17 | [Surogate Rune 26B-A4B v3 (RTX PRO 6000)](/models/surogate-rune-26b-a4b-v3-rtxpro6000-a2) | 66.47 | 69.74 | 88.30 | 85.97 | 48.97 | 65.43–67.01 | Not established |
| 18 | [reflex 4B (kshetrajna12)](/models/reflex-4b) | 65.24 | 51.49 | 86.84 | 68.77 | 63.14 | 58.56–66.35 | Statistical tie with jev-local Qwen3.5-9B |
| 19 | [jev-local (Qwen3.5-9B)](/models/jev-local) | 65.24 | 56.28 | 77.82 | 72.98 | 58.86 | 63.46–66.50 | Statistical tie with djev |
| 20 | [djev (Maisa, diffusion-gemma)](/models/djev) | 64.16 | 72.33 | 80.41 | 91.00 | 48.22 | 63.07–65.00 | Statistical tie with Raw Qwen3 4B Instruct 2507 direct logits |
| 21 | [Raw Qwen3 4B Instruct 2507 direct logits](/models/raw-qwen3-4b-instruct-2507) | 62.14 | 54.06 | 52.95 | 89.23 | 63.36 | 59.35–64.22 | Statistical tie with jqv Qwen3-32B |
| 22 | [jqv (Qwen3-32B zero-shot)](/models/jqv) | 60.77 | 49.09 | 86.77 | 83.36 | 51.17 | 51.25–64.26 | Statistical tie with JevK5 v0.2.0 |
| 23 | [JevK5 v0.2.0](/models/jevk5-v0-2-0) | 58.11 | 46.70 | 84.87 | 90.86 | 63.07 | 47.56–68.24 | Statistical tie with Qwen3.5-9B Jev-like data-mix v2 |
| 24 | [Qwen3.5-9B Jev-like data-mix v2](/models/qwen35-9b-jev-data-mix-v2) | 53.02 | 60.45 | 80.86 | 82.00 | 45.69 | 51.88–53.88 | Published separation with Standard One 8B |
| 25 | [Standard One 8B (Standard Thinking)](/models/standardone-8b) | 47.79 | 59.60 | 83.16 | 92.48 | 43.28 | 46.67–48.54 | Statistical tie with NInfer Qwen3.8-Flash-Next mixed |
| 26 | [NInfer Qwen3.8-Flash-Next mixed](/models/ninfer-qwen3-8-flash-next) | 47.47 | 67.20 | 88.54 | 88.61 | 42.53 | 46.70–47.88 | Not established |
| 27 | [Instinct Dual 4B](/models/instinct-dual-4b) | 46.97 | 43.12 | 88.30 | 81.95 | 60.19 | 37.86–56.77 | Not established |
| 28 | [swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4)](/models/swanone) | 46.56 | 71.22 | 87.06 | 84.74 | 42.15 | 45.81–46.94 | Statistical tie with Raw Qwen3 8B direct logits |
| 29 | [Raw Qwen3 8B direct logits](/models/raw-qwen3-8b) | 45.22 | 51.14 | 49.19 | 86.84 | 45.54 | 36.89–46.74 | Statistical tie with decider-2b |
| 30 | [decider-2b (Mapika)](/models/decider-2b) | 45.10 | 42.34 | 71.53 | 94.40 | 64.92 | 33.82–54.26 | Statistical tie with system-one Qwen3-8B |
| 31 | [system-one (Qwen3-8B, Sean Goedecke)](/models/system-one-sg) | 44.14 | 50.47 | 49.38 | 90.88 | 44.97 | 36.79–45.66 | Statistical tie with system-one-open |
| 32 | [system-one-open (Gemma 4 E2B LoRA on an L4)](/models/system-one-open) | 42.44 | 41.64 | 71.63 | 78.27 | 68.33 | 33.84–51.92 | Statistical tie with Autoloops Gemma 4 31B IT |
| 33 | [Autoloops – Gemma 4 31B IT](/models/autoloops-gemma-4-31b-it) | 40.53 | 76.68 | 85.84 | 83.92 | 39.59 | 40.09–40.76 | Statistical tie with GPT-6 Luna (low) |
| 34 | [GPT-6 Luna (low reasoning effort)](/models/gpt-6-luna-low) | 40.48 | 95.29 | 94.92 | 73.20 | 39.06 | 40.27–40.66 | Published separation with GPT-6 Luna (medium) |
| 35 | [GPT-6 Luna (default medium reasoning effort)](/models/gpt-6-luna-medium) | 38.75 | 96.24 | 95.56 | 73.19 | 38.32 | 38.57–38.93 | Statistical tie with JevOne |
| 36 | [JevOne](/models/jevone) | 38.25 | 54.13 | 84.94 | 90.43 | 39.84 | 37.38–38.79 | Statistical tie with kev 4B |
| 37 | [kev 4B (research preview)](/models/kev-4b) | 38.07 | 39.86 | 67.64 | 85.48 | 65.78 | 27.48–46.77 | Statistical tie with kev 8B |
| 38 | [kev 8B (research preview)](/models/kev-8b) | 34.15 | 48.30 | 71.02 | 84.14 | 40.42 | 26.79–37.48 | Statistical tie with open-alternative-jev Qwen3.5-4B |
| 39 | [open-alternative-jev (Qwen3.5-4B, IkerMoel)](/models/open-alternative-jev) | 33.56 | 37.36 | 76.87 | 91.29 | 63.28 | 25.75–41.66 | Statistical tie with Bespoke Nimble 9B |
| 40 | [Bespoke Nimble 9B (Bespoke Labs)](/models/nimble-9b) | 31.82 | 63.75 | 77.16 | 82.95 | 36.75 | 31.19–32.27 | Statistical tie with Malkuth-2B |
| 41 | [Malkuth-2B (newfull5, Kev post-train)](/models/malkuth-2b) | 29.90 | 35.54 | 75.19 | 91.80 | 65.57 | 20.96–38.98 | Statistical tie with openjev-sglang Qwen3.6-35B-A3B |
| 42 | [openjev-sglang (Qwen3.6-35B-A3B on SGLang)](/models/openjev-sglang) | 29.03 | 58.65 | 82.87 | 78.07 | 35.63 | 28.42–29.40 | Published separation with decider-35b-a3b |
| 43 | [decider-35b-a3b (Mapika)](/models/decider-35b-a3b) | 27.50 | 60.46 | 82.04 | 91.00 | 34.39 | 26.93–27.87 | Statistical tie with local-jev Qwen3.5-4B |
| 44 | [local-jev Qwen3.5-4B](/models/localjev-qwen3-5-4b) | 25.82 | 33.75 | 82.45 | 83.74 | 59.27 | 19.89–32.96 | Not established |
| 45 | [Nemotron Diffusion 8B (pst2154, optimized vLLM)](/models/nemotron-diffusion-8b) | 25.68 | 33.84 | 76.25 | 93.74 | 55.49 | 18.60–32.78 | Not established |
| 46 | [Open-Jev 9B (Zefan Cai)](/models/open-jev-zefan-9b) | 24.36 | 63.82 | 81.53 | 73.59 | 33.05 | 23.93–24.64 | Statistical tie with Decision 2B v59 |
| 47 | [Decision 2B (FlyMy.AI, v59)](/models/decision-2b) | 22.52 | 31.31 | 86.22 | 90.36 | 66.37 | 16.68–28.77 | Statistical tie with GPT-5.6 Luna (low) |
| 48 | [GPT-5.6 Luna (low reasoning effort)](/models/gpt-5-6-luna-low) | 22.37 | 94.33 | 94.71 | 74.06 | 30.67 | 22.26–22.46 | Statistical tie with typecastlm |
| 49 | [typecastlm (Mikhail Gribov, Qwen3.5-4B computed head)](/models/typecastlm) | 21.83 | 31.24 | 76.71 | 92.04 | 63.96 | 16.23–28.58 | Statistical tie with JEV Qwen3.5-9B Base NVFP4 |
| 50 | [JEV Qwen3.5-9B Base NVFP4](/models/jev-qwen3-5-9b-base-nvfp4) | 20.07 | 32.24 | 80.93 | 93.54 | 47.59 | 15.19–25.51 | Statistical tie with Gemini 3.1 Flash-Lite |
| 51 | [Gemini 3.1 Flash-Lite](/models/gemini-3-1-flash-lite) | 19.58 | 77.59 | 74.68 | 80.05 | 29.76 | 19.30–19.83 | Not established |
| 52 | [AutoJev-27B (denis-pplx, Qwen3.8-27B)](/models/autojev-27b) | 19.54 | 72.78 | 87.70 | 87.62 | 29.37 | 19.30–19.63 | Not established |
| 53 | [AutoJev-27B (RTX PRO 6000)](/models/autojev-27b-rtxpro6000-a2) | 19.48 | 72.75 | 86.68 | 87.07 | 29.37 | 19.26–19.60 | Not established |
| 54 | [NInfer Qwen3.8-27B NVFP4](/models/ninfer-qwen3-8-27b) | 18.74 | 65.47 | 85.91 | 89.86 | 29.12 | 18.45–18.92 | Not established |
| 55 | [Eikos-27B (caiovicentino1, Qwen3.8-27B)](/models/eikos-27b) | 18.50 | 75.08 | 86.29 | 87.63 | 28.69 | 18.28–18.60 | Not established |
| 56 | [NInfer Qwen3.8-27B NVFP4 (T=1.5)](/models/ninfer-qwen3-8-27b-t1-5) | 18.48 | 61.08 | 86.49 | 89.86 | 29.12 | 18.18–18.65 | Statistical tie with Instinct Qwen3.8-27B |
| 57 | [Instinct (ZooWork, Qwen3.8-27B)](/models/instinct) | 18.32 | 62.75 | 84.96 | 81.68 | 29.16 | 18.03–18.51 | Published separation with OpenJev (thinking, BF16) |
| 58 | [OpenJev (thinking, BF16)](/models/openjev-thinking) | 17.94 | 84.21 | 83.10 | 73.86 | 28.51 | 17.76–18.08 | Published separation with djev (thinking) |
| 59 | [djev (thinking)](/models/djev-thinking) | 17.40 | 77.35 | 95.66 | 72.32 | 28.13 | 17.28–17.51 | Published separation with LitJev Qwen3.8-27B |
| 60 | [LitJev (Qwen3.8-27B)](/models/litjev) | 16.31 | 58.32 | 84.45 | 68.22 | 28.36 | 16.03–16.48 | Not established |
| 61 | [Bev / Bonsai 27B](/models/bev-bonsai-27b) | 15.77 | 53.08 | 77.68 | 72.73 | 28.23 | 15.13–16.05 | Not established |
| 62 | [Raw Phi-4 mini direct logits](/models/raw-phi-4-mini) | 15.22 | 27.62 | 71.41 | 89.17 | 53.22 | 10.26–20.74 | Statistical tie with OpenSourceJev Qwen3.5-4B Q4_K_M |
| 63 | [OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)](/models/opensourcejev-qwen35-4b-q4km) | 13.25 | 25.77 | 76.04 | 72.51 | 69.14 | 8.98–18.31 | Statistical tie with reflex-27b |
| 64 | [reflex-27b (Qwen3.8-27B)](/models/reflex-27b) | 13.19 | 62.79 | 85.93 | 69.20 | 25.80 | 12.99–13.30 | Published separation with Open-Jev 2B |
| 65 | [Open-Jev 2B (Zefan Cai)](/models/open-jev-zefan-2b) | 9.07 | 33.56 | 73.62 | 75.76 | 33.05 | 6.62–11.54 | Statistical tie with GLiNER2 large |
| 66 | [GLiNER2 large (Fastino)](/models/gliner2-large) | 8.40 | 22.49 | 42.38 | 64.78 | 77.59 | 5.21–12.34 | Statistical tie with Qwen3-Reranker-4B |
| 67 | [Qwen3-Reranker-4B](/models/qwen3-reranker-4b) | 7.06 | 21.03 | 76.26 | 79.61 | 48.40 | 4.26–10.74 | Statistical tie with DeepSeek V4.1 Flash |
| 68 | [DeepSeek V4.1 Flash (thinking default)](/models/deepseek-v4-1-flash) | 6.65 | 93.69 | 96.92 | 69.40 | 19.09 | 6.62–6.66 | Statistical tie with SimpleJev Qwen3.5-0.8B |
| 69 | [SimpleJev (Qwen3.5-0.8B, CPU)](/models/simplejev-qwen3-5-0-8b) | 4.00 | 16.75 | 46.67 | 59.07 | 70.31 | 2.06–6.66 | Statistical tie with SimpleJev Qwen3.8-27B |
| 70 | [SimpleJev Qwen3.8-27B](/models/simplejev-qwen3-8-27b) | 3.36 | 72.75 | 87.20 | 74.92 | 14.90 | 3.33–3.37 | Statistical tie with decision-machine-1 |
| 71 | [decision-machine-1 (milliseconds.ai)](/models/decision-machine-1) | 3.24 | 14.82 | 81.18 | 92.77 | 56.30 | 1.73–5.39 | Statistical tie with GLiNER2.5 multi |
| 72 | [GLiNER2.5 multi (Fastino, 287M)](/models/gliner2-5-multi) | 2.72 | 14.00 | 57.96 | 66.57 | 86.62 | 1.20–5.10 | Not established |
| 73 | [Bosun v3.1 0.6B](/models/bosun-v31-0-6b) | 2.50 | 13.58 | 65.16 | 61.76 | 77.55 | 1.15–4.56 | Not established |
| 74 | [GLiNER2 (Fastino, gliner2.5-base)](/models/gliner2-5-base) | 2.27 | 13.48 | 35.57 | 70.33 | 86.62 | 0.90–4.50 | Not established |
| 75 | [Deem 0.8B v1](/models/deem-0-8-v1) | 2.14 | 13.02 | 38.88 | 84.32 | 80.29 | 0.90–4.34 | Not established |
| 76 | [JevAct (einptein, jev1-2b-v2)](/models/jevact) | 1.52 | 11.24 | 62.54 | 76.43 | 68.56 | 0.52–3.18 | Statistical tie with CLM-8B |
| 77 | [CLM-8B (Contrastive-LM, clm-latest)](/models/clm-8b) | 1.51 | 11.43 | 48.48 | 93.29 | 50.50 | 0.48–3.17 | Statistical tie with kev 0.6B |
| 78 | [kev 0.6B (research preview)](/models/kev-0-6b) | 1.26 | 10.34 | 67.65 | 87.22 | 80.09 | 0.49–2.50 | Statistical tie with Raw Qwen3 0.6B direct logits |
| 79 | [Raw Qwen3 0.6B direct logits](/models/raw-qwen3-0-6b) | 1.08 | 10.56 | 21.45 | 90.39 | 77.58 | 0.27–2.61 | Statistical tie with GLiNER2.5 small |
| 80 | [GLiNER2.5 small (Fastino, 74M)](/models/gliner2-5-small) | 0.89 | 9.19 | 55.90 | 77.36 | 86.62 | 0.23–2.20 | Statistical tie with Raw Qwen3 1.7B direct logits |
| 81 | [Raw Qwen3 1.7B direct logits](/models/raw-qwen3-1-7b) | 0.85 | 9.67 | 21.57 | 90.22 | 68.55 | 0.15–2.28 | Statistical tie with Mirror |
| 82 | [Mirror](/models/mirror) | 0.22 | 5.60 | 43.21 | 63.64 | 89.27 | 0.01–0.90 | Statistical tie with ZeroEntropy zerank-2 |
| 83 | [ZeroEntropy zerank-2](/models/zerank-2) | 0.13 | 4.76 | 81.71 | 80.28 | 48.40 | 0.01–0.49 | Statistical tie with jeff |
| 84 | [jeff (Logan Markewich, GLiFormer 400M)](/models/jeff) | 0.12 | 4.43 | 80.27 | 55.76 | 81.09 | 0.00–0.51 | Statistical tie with smalljev semantic-v9 |
| 85 | [smalljev semantic-v9](/models/smalljev) | 0.07 | 3.66 | 73.09 | 86.46 | 60.77 | 0.00–0.38 | Not established |
| 86 | [Laya multilingual](/models/laya-multilingual) | 0.02 | 2.35 | 43.58 | 73.72 | 82.16 | 0.00–0.27 | Not established |
| 87 | [OpenDecision (ModernBERT-large zero-shot)](/models/opendecision) | 0.01 | 2.07 | 72.52 | 86.58 | 79.11 | 0.00–0.19 | Statistical tie with BAAI bge-reranker-v2-m3 |
| 88 | [BAAI bge-reranker-v2-m3](/models/bge-reranker-v2-m3) | 0.00 | 0.00 | 83.33 | 90.55 | 59.33 | 0.00–0.00 | Statistical tie with Certo v1 |
| 89 | [Certo v1 (AltSlate Labs)](/models/certo) | 0.00 | 0.00 | 87.52 | 91.43 | 96.91 | 0.00–0.00 | Statistical tie with Decision Fast v53a |
| 90 | [Decision Fast (FlyMy.AI, v53a)](/models/decision-fast) | 0.00 | 0.00 | 76.08 | 91.44 | 80.09 | 0.00–0.00 | Statistical tie with Alibaba GTE Reranker ModernBERT-base |
| 91 | [Alibaba GTE Reranker ModernBERT-base](/models/gte-reranker-modernbert-base) | 0.00 | 0.00 | 75.20 | 91.49 | 69.27 | 0.00–0.00 | Statistical tie with kev 0.5B |
| 92 | [kev 0.5B](/models/kev-0-5b) | 0.00 | 0.00 | 64.98 | 88.45 | 80.09 | 0.00–0.00 | Statistical tie with Laya |
| 93 | [Laya (Convai Innovations, ModernBERT-large 421M)](/models/laya) | 0.00 | 0.00 | 73.71 | 73.89 | 84.89 | 0.00–0.00 | Statistical tie with lev-350m |
| 94 | [lev-350m (Franck Verrot, LFM2.5-350M)](/models/lev-350m) | 0.00 | 0.00 | 77.68 | 93.63 | 80.04 | 0.00–0.00 | Statistical tie with Qwen3.5-0.8B Decision Model |
| 95 | [Qwen3.5-0.8B Decision Model (Mourad Ghafiri)](/models/mghafiri-qwen3-5-0-8b-decision-model) | 0.00 | 0.00 | 73.57 | 71.82 | 79.52 | 0.00–0.02 | Statistical tie with Mixedbread mxbai-rerank-base-v2 |
| 96 | [Mixedbread mxbai-rerank-base-v2](/models/mxbai-rerank-base-v2) | 0.00 | 0.00 | 86.61 | 89.34 | 60.34 | 0.00–0.00 | Statistical tie with Needle 3 (2-bit) |
| 97 | [Needle 3 (Cactus, 2-bit, local CPU)](/models/needle-3) | 0.00 | 0.00 | 0.00 | 34.13 | 61.57 | 0.00–0.00 | Statistical tie with Needle 3 (options as tools) |
| 98 | [Needle 3, options as tools (post-hoc adapter mode)](/models/needle-3-tools) | 0.00 | 0.00 | 0.00 | 41.00 | 61.57 | 0.00–0.00 | Statistical tie with open-jev-deberta-v3-large |
| 99 | [open-jev-deberta-v3-large (local CPU)](/models/open-jev-deberta-v3-large) | 0.00 | 0.00 | 77.06 | 68.25 | 77.59 | 0.00–0.00 | Statistical tie with Open Jev JSON Canvas |
| 100 | [Open Jev JSON Canvas (JoshuaSP)](/models/open-jev-json-canvas-joshuasp) | 0.00 | 77.14 | 0.00 | 85.57 | 49.30 | 0.00–0.00 | Statistical tie with openJev Verdict |
| 101 | [openJev Verdict (heman10x, ModernBERT-base 151M)](/models/openjev-verdict) | 0.00 | 0.00 | 52.22 | 83.88 | 86.62 | 0.00–0.01 | Statistical tie with openJev Verdict 1.4 |
| 102 | [openJev Verdict 1.4](/models/openjev-verdict-1-4) | 0.00 | 0.00 | 80.33 | 80.64 | 86.62 | 0.00–0.00 | Statistical tie with Qwen3.8-27B (Chutes TEE) |
| 103 | [Qwen3.8 27B (Chutes TEE)](/models/qwen3-8-27b-chutes) | 0.00 | 95.56 | 98.09 | 56.85 | 0.00 | 0.00–0.00 | Statistical tie with verdict-small |
| 104 | [verdict-small (Manavarya09, multilingual-e5-small 118M)](/models/verdict-small) | 0.00 | 0.00 | 59.45 | 82.06 | 100.00 | 0.00–0.01 | Statistical tie with Von 395M |
| 105 | [Von (wfzyx, Option-Marker 395M)](/models/von-395m) | 0.00 | 0.00 | 83.48 | 75.72 | 82.71 | 0.00–0.00 | Not established |
| 106 | [Laya typed-decisions](/models/laya-typed-decisions) | 0.00 | 0.00 | 83.34 | 63.38 | 82.53 | 0.00–0.00 | Not established |
| Unranked | [classifier.dev (fast tier)](/models/classifier-dev-fast) | 74.66 | 75.81 | 89.19 | 81.99 | 58.89 | 73.50–75.20 | Not established |
| Unranked | [SimpleJev Qwen3.6-35B-A3B](/models/simplejev-qwen3-6-35b-a3b) | 0.00 | 0.00 | 78.50 | 75.18 | 35.15 | 0.00–0.00 | Not established |
| Unranked | [Decision-4B (Eval Engine / Chromia)](/models/evalengine-decision-4b) | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not established |
| Unranked | [Jobe Qwen3.5-4B (frozen)](/models/jobe-qwen3-5-4b) | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not established |
| Unranked | [mica-v01-4b](/models/mica-v0-1-4b) | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not established |
| Unranked | [OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)](/models/openjev-razorback16) | Not measured | Not measured | Not measured | Not measured | Not measured | Not measured | Not established |

[JevBench method and result notes](/benchmarks/jevbench) · [Pinned official results](https://benchmarkheaven.com/api/jevbench/v1.5.4)

JevBench by [Florian Standhartinger](https://github.com/fstandhartinger) and contributors. [MIT license and copyright notice](/licenses/jevbench-mit.txt).

## Benchmarks with a different scope

[DecisionBench by Hanno Labs](https://hannolabs.ai/field-notes/decisionbench). A separate benchmark for typed decisions. Its applied suite and reasoning track stay separate; results are linked at the source rather than imported into this table.

## Decision task benchmarks

Compare scores only when the sample, label mapping, language, and evaluation setup match. A subset or binary conversion keeps its own result.

Original benchmark coverage: 40 source tables across seven benchmarks. The linked pages include every imported numeric metric and each identified evaluated configuration, with JSON downloads. Anonymous submissions and conflicting source values are labeled.

- [WinoGrande](/benchmarks/winogrande): Resolve an ambiguous reference in a sentence.
- [FinancialPhraseBank](/benchmarks/financialphrasebank): Classify financial sentiment as positive, negative, or neutral.
- [RAGTruth](/benchmarks/ragtruth): Detect unsupported claims against retrieved evidence.
- [JudgeBench](/benchmarks/judgebench): Choose the objectively better response from a pair.
- [BBH](/benchmarks/bbh): Solve challenging reasoning tasks from BIG-Bench.
- [JevBench](/benchmarks/jevbench): Evaluate decision systems; public-subset accuracy is separate from the full composite.
- [TabFact](/benchmarks/tabfact): Check whether a table supports or refutes a statement.
- [ContractNLI](/benchmarks/contractnli): Classify a contract hypothesis and identify supporting evidence.
- [Circa](/benchmarks/circa): Interpret an indirect answer to a yes/no question.
- [Belebele](/benchmarks/belebele): Answer passage-comprehension questions across language variants.
- [TruthfulQA](/benchmarks/truthfulqa): Test resistance to common misconceptions; a binary conversion is a separate setup.

## Choose a system for your own decisions

### Match the decision to the test

Check the rubric, input length, answer options, and model configuration. A reranker adapted to pick an option and a native decision model can answer the same question through different serving paths.

### Check confidence against errors

Calibration asks whether reported probabilities match observed outcomes. Test an escalation threshold on your own labeled cases before letting a high-confidence answer approve an action or bypass a larger model.

### Price the whole decision

Compare cost per completed decision alongside latency and accuracy. JevBench estimates some hosting costs from base-model reference prices and adjusts self-hosted latency. Those assumptions can change which system fits your budget.

## Questions

### What does Perplexity Decider’s 85.71% measure?

The 85.71% is Perplexity’s reported accuracy across a fixed 7,210-row panel of 11 tasks from September 2026. Task samples have different sizes, and Decider was measured through the Perplexity API. The percentage measures correctness on those samples; it does not establish probability calibration or replace JevBench’s full composite.

### What does Perplexity Decider cost?

Perplexity lists pplx-decider-v1-27b at $0.04 per million input tokens, with free output tokens and no per-request fee. State, images, and all questions count toward input usage. Use the returned input-token count to price a completed decision; a short label alone does not determine the request cost.

### How are decision models different from reasoning models?

Decision models answer within a defined set of choices or score levels, often with probabilities your code can use. Reasoning models are evaluated on solving broader problems. A general-purpose model can act as a decision system, but its result depends on the prompt, adapter, and serving setup.

### Which benchmarks evaluate AI decision models?

JevBench evaluates typed decision systems across intelligence, calibration, speed, and cost. DecisionBench from Hanno Labs publishes task-level accuracy and probability-quality measures, with applied decisions and reasoning in separate tracks. This page mirrors the pinned JevBench release and links to DecisionBench directly; the two scores are not interchangeable.

### Does the Decision Models category change overall rankings?

The category groups decision-system evidence for discovery. JevBench retains its official ranking and remains outside overall and category score calculations because its composite includes serving latency, pricing assumptions, and adapters. Model pages keep the evaluated configuration separate from its base model, so scores are not inherited between them.

## Related benchmarks

[Benchmark directory](/benchmarks) · [Reasoning models](/reasoning) · [Agentic models](/agentic) · [How the rankings work](/methodology)
