Cloudflare Decision Index 0.2.1: API-Bank
We show this table for reference; we do not rank on it.
Published Cloudflare and community decision-model configurations.
Evaluation results published by Cloudflare and upstream community contributors.
accuracy on Cloudflare Decision Index 0.2.1: API-Bank — October 1, 2026 capture
We mirror the published accuracy view for Cloudflare Decision Index 0.2.1: API-Bank. Clef-flash leads the public snapshot at 93.1%, followed by Clef (91.9%) and Jev (88.2%). We do not use these results to rank models overall.
Clef-flash
Cloudflare
Cloudflare self-reported; Routing head + LoRA; 36/38 index tasks; median 38.8 ms; p95 122.4 ms
Clef
Cloudflare
Cloudflare self-reported; Routing head + LoRA; 36/38 index tasks; median 209.3 ms; p95 238.6 ms
Jev
TypeSafe
Upstream community result; Hosted API (closed); 38/38 index tasks; median 524.1 ms; p95 536 ms
64 configurationsDecision ModelsDated mixed-source configurationsDisplay onlyUpdated October 1, 2026 capture
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Clef-flashCloudflareCloudflare self-reported; Routing head + LoRA; 36/38 index tasks; median 38.8 ms; p95 122.4 ms | 93.1% | Not reported | Unknown |
| 2 | ClefCloudflareCloudflare self-reported; Routing head + LoRA; 36/38 index tasks; median 209.3 ms; p95 238.6 ms | 91.9% | Not reported | Unknown |
| 3 | JevTypeSafeUpstream community result; Hosted API (closed); 38/38 index tasks; median 524.1 ms; p95 536 ms | 88.2% | Not reported | Unknown |
| 4 | Eikos-27B [FP8]Creator not reportedUpstream community result; LoRA; 38/38 index tasks; median 129.7 ms; p95 509.8 ms | 86.0% | Not reported | Unknown |
| 5 | simple-jev · Qwen3.8-27B (featherless)Creator not reportedUpstream community result; inference technique; 38/38 index tasks; median 373 ms; p95 2796.8 ms | 85.4% | Not reported | Unknown |
| 6 | Decider chat · Gemma-4-31BCreator not reportedUpstream community result; inference technique; 38/38 index tasks; median 108.5 ms; p95 1114.7 ms | 85.0% | Not reported | Unknown |
| 7 | reflex 27B [FP8 · wide choice]Creator not reportedUpstream community result; inference technique; 38/38 index tasks; median 108.3 ms; p95 1239.7 ms | 85.0% | Not reported | Unknown |
| 8 | AutoJev-27BCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 101.4 ms; p95 1052.5 ms | 84.1% | Not reported | Unknown |
| 9 | Jebadiah 27BCreator not reportedUpstream community result; LoRA; 38/38 index tasks; median 110.4 ms; p95 591.1 ms | 84.1% | Not reported | Unknown |
| 10 | JevfireCreator not reportedUpstream community result; inference technique; 38/38 index tasks; median 89.6 ms; p95 447.5 ms | 84.1% | Not reported | Unknown |
| 11 | JoshuaSP diffusiongemma (open-jev)Creator not reportedUpstream community result; inference technique; 38/38 index tasks; median 260.4 ms; p95 537.3 ms | 83.7% | Not reported | Unknown |
| 12 | Surogate Rune 26B-A4B v3 [bf16]Creator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 120.5 ms; p95 682.8 ms | 83.3% | Not reported | Unknown |
| 13 | Bespoke Nimble 9B v2Creator not reportedUpstream community result; LoRA; 38/38 index tasks; median 77.3 ms; p95 621.4 ms | 80.5% | Not reported | Unknown |
| 14 | Decider chat · Qwen3.6-27BCreator not reportedUpstream community result; inference technique; 38/38 index tasks; median 83.6 ms; p95 965.6 ms | 79.7% | Not reported | Unknown |
| 15 | levCreator not reportedUpstream community result; LoRA + head; 38/38 index tasks; median 70.8 ms; p95 687.7 ms | 79.1% | Not reported | Unknown |
| 16 | JPT-9BCreator not reportedUpstream community result; LoRA; 38/38 index tasks; median 148.6 ms; p95 461.9 ms | 78.7% | Not reported | Unknown |
| 17 | Decision 1.0 NoxCreator not reportedUpstream community result; head / adapter; 38/38 index tasks; median 53.2 ms; p95 242.1 ms | 77.6% | Not reported | Unknown |
| 18 | Intern-Decision-4BCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 44.2 ms; p95 115.4 ms | 76.0% | Not reported | Unknown |
| 19 | Hopper (G) 1.2Creator not reportedUpstream community result; LoRA; 38/38 index tasks; median 23.3 ms; p95 238.2 ms | 75.6% | Not reported | Unknown |
| 20 | Decision 1.0 LuxCreator not reportedUpstream community result; head / adapter; 38/38 index tasks; median 51 ms; p95 347.5 ms | 74.0% | Not reported | Unknown |
| 21 | Winnow-E4B [Q8_0]Creator not reportedUpstream community result; LoRA; 38/38 index tasks; median 45 ms; p95 188.3 ms | 73.0% | Not reported | Unknown |
| 22 | djevCreator not reportedUpstream community result; inference technique; 38/38 index tasks; median 84.4 ms; p95 211.2 ms | 70.7% | Not reported | Unknown |
| 23 | JPT-4BCreator not reportedUpstream community result; LoRA; 38/38 index tasks; median 139.1 ms; p95 405.7 ms | 68.7% | Not reported | Unknown |
| 24 | JevK5Creator not reportedUpstream community result; LoRA; 38/38 index tasks; median 22 ms; p95 226.3 ms | 68.5% | Not reported | Unknown |
| 25 | Winnow-12B [Q8_0]Creator not reportedUpstream community result; LoRA; 38/38 index tasks; median 72.5 ms; p95 356.2 ms | 67.3% | Not reported | Unknown |
| 26 | NeoHorse-Jev-4BCreator not reportedUpstream community result; head / adapter; 38/38 index tasks; median 49 ms; p95 273.9 ms | 66.1% | Not reported | Unknown |
| 27 | Decision 1.0 SolCreator not reportedUpstream community result; head / adapter; 38/38 index tasks; median 38.1 ms; p95 101 ms | 62.0% | Not reported | Unknown |
| 28 | Bosun v3.1 1.7BCreator not reportedUpstream community result; LoRA + head; 38/38 index tasks; median 40.3 ms; p95 173.5 ms | 60.8% | Not reported | Unknown |
| 29 | Decider 4BCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 12.6 ms; p95 188.5 ms | 60.4% | Not reported | Unknown |
| 30 | razorback16 openjev diffusiongemma (NVFP4, vLLM) [one read, NVFP4]Creator not reportedUpstream community result; inference technique; 38/38 index tasks; median 37.7 ms; p95 117.8 ms | 57.3% | Not reported | Unknown |
| 31 | Decider 35B-A3B [NVFP4]Creator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 101.4 ms; p95 210.9 ms | 56.5% | Not reported | Unknown |
| 32 | Kev 9BCreator not reportedUpstream community result; LoRA + head; 38/38 index tasks; median 51.4 ms; p95 187.9 ms | 56.3% | Not reported | Unknown |
| 33 | Jet v6.2Creator not reportedUpstream community result; LoRA; 38/38 index tasks; median 46.8 ms; p95 371.1 ms | 54.3% | Not reported | Unknown |
| 34 | Bosun v3.1 0.6BCreator not reportedUpstream community result; LoRA + head; 38/38 index tasks; median 38.7 ms; p95 107.3 ms | 53.1% | Not reported | Unknown |
| 35 | Jev-OmniCreator not reportedUpstream community result; LoRA + head; 38/38 index tasks; median 54.9 ms; p95 627.9 ms | 52.4% | Not reported | Unknown |
| 36 | this-that 1.2Creator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 44.3 ms; p95 57.2 ms | 52.4% | Not reported | Unknown |
| 37 | Kev 4BCreator not reportedUpstream community result; LoRA + head; 38/38 index tasks; median 52.1 ms; p95 141.1 ms | 51.2% | Not reported | Unknown |
| 38 | open-jev (pngwn)Creator not reportedUpstream community result; LoRA; 38/38 index tasks; median 112.3 ms; p95 724.1 ms | 48.4% | Not reported | Unknown |
| 39 | JobeCreator not reportedUpstream community result; inference technique; 38/38 index tasks; median 53.5 ms; p95 2910.3 ms | 46.1% | Not reported | Unknown |
| 40 | openvonsCreator not reportedUpstream community result; inference technique; 38/38 index tasks; median 21.2 ms; p95 113.3 ms | 43.5% | Not reported | Unknown |
| 41 | Kev 0.8BCreator not reportedUpstream community result; LoRA + head; 38/38 index tasks; median 41.6 ms; p95 95.9 ms | 42.9% | Not reported | Unknown |
| 42 | Decision 1.0 EosCreator not reportedUpstream community result; head / adapter; 38/38 index tasks; median 39.7 ms; p95 84.3 ms | 36.6% | Not reported | Unknown |
| 43 | Decider 2B [FP8]Creator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 8.1 ms; p95 134.9 ms | 31.7% | Not reported | Unknown |
| 44 | JPT-0.8BCreator not reportedUpstream community result; LoRA; 38/38 index tasks; median 117.5 ms; p95 341.2 ms | 30.1% | Not reported | Unknown |
| 45 | MoJevCreator not reportedUpstream community result; head / adapter; 38/38 index tasks; median 46.4 ms; p95 55.1 ms | 29.1% | Not reported | Unknown |
| 46 | Intern-Decision-2BCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 33.5 ms; p95 58 ms | 28.5% | Not reported | Unknown |
| 47 | LavoirCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 20 ms; p95 49.5 ms | 28.3% | Not reported | Unknown |
| 48 | Intern-Decision-0.8BCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 33.7 ms; p95 39.2 ms | 21.1% | Not reported | Unknown |
| 49 | Julia 1Creator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 5.8 ms; p95 15.3 ms | 18.3% | Not reported | Unknown |
| 50 | Decision 1.0 KaiCreator not reportedUpstream community result; head / adapter; 38/38 index tasks; median 30.4 ms; p95 106.8 ms | 15.9% | Not reported | Unknown |
| 51 | GLiNER 2.5 baseCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 14.1 ms; p95 62.1 ms | 15.6% | Not reported | Unknown |
| 52 | Decision 1.0 LexCreator not reportedUpstream community result; head / adapter; 38/38 index tasks; median 29.8 ms; p95 106 ms | 14.0% | Not reported | Unknown |
| 53 | LayaCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 5.8 ms; p95 222.5 ms | 11.5% | Not reported | Unknown |
| 54 | Lumma-Fev-0.1BCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 20.9 ms; p95 112.5 ms | 8.3% | Not reported | Unknown |
| 55 | GLiNER2.5-DecideCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 23.3 ms; p95 228.6 ms | 7.7% | Not reported | Unknown |
| 56 | CLM-v0.1-8BCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 46.8 ms; p95 130.9 ms | 6.1% | Not reported | Unknown |
| 57 | Qwen-2.5-1B-RLCDCreator not reportedUpstream community result; inference technique; 38/38 index tasks; median 43.8 ms; p95 75.2 ms | 4.3% | Not reported | Unknown |
| 58 | Lumma-Fev-0.6BCreator not reportedUpstream community result; LoRA; 38/38 index tasks; median 21.3 ms; p95 219.7 ms | 4.1% | Not reported | Unknown |
| 59 | jeffCreator not reportedUpstream community result; inference technique; 38/38 index tasks; median 21.8 ms; p95 67.6 ms | 3.1% | Not reported | Unknown |
| 60 | LFM2.5-350M-RLCDCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 26.7 ms; p95 174.3 ms | 2.4% | Not reported | Unknown |
| 61 | GLiNER 2.5 smallCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 13.6 ms; p95 35.2 ms | 2.0% | Not reported | Unknown |
| 62 | LFM2.5-2.6B-RLCDCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 39.3 ms; p95 95.4 ms | 1.8% | Not reported | Unknown |
| 63 | GLiNER 2.5 multilingualCreator not reportedUpstream community result; full fine-tune; 38/38 index tasks; median 14 ms; p95 92.4 ms | 1.6% | Not reported | Unknown |
| 64 | system-one-gemmaCreator not reportedUpstream community result; LoRA + head; 38/38 index tasks; median 31.9 ms; p95 748 ms | 1.2% | Not reported | Unknown |
How to read this source comparison
The source combines Cloudflare self-reported Clef runs with upstream community configurations. Each row retains its source engine and evidence label; model names alone do not establish equivalent serving settings.
The 0.2.1 composite uses 38 panel benchmarks grouped into five weighted areas. It applies chance correction; ForecastBench uses its Brier transformation. Additional task views retain the source metric and panel-inclusion metadata. The snapshot lists 48 task metrics, while the launch refers to 43 evaluations including latency rows.
Clef and Clef-Flash each cover 36 of 38 index tasks; HLE and iSarcasmEval are missing. Missing values are not inserted as zero. The composite is retained as published without recomputing it or filling gaps.
Community results use the upstream September 28 snapshot, which declares one NVIDIA RTX PRO 6000. Cloudflare run hardware is not established by this artifact. Latency comparisons across these configurations do not isolate model speed. The October 1 date marks capture, not evaluation.
These views remain separate from JevBench and canonical provider benchmark results. They do not enter overall or category rankings. Equal scores do not establish statistical ties.
The published Cloudflare Decision Index 0.2.1: API-Bank snapshot places Clef-flash first at 93.1%. The third row is 4.9 points behind. The broader top-10 range is 9.0 points, so many of the published results sit in a relatively narrow band.
64 configurations are shown for Cloudflare Decision Index 0.2.1: API-Bank. The benchmark falls in the Decision Models category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Cloudflare Decision Index 0.2.1: API-Bank is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
Freshness and provenance
Version
Cloudflare comparison on Decision Index 0.2.1
Refresh cadence
Reviewed dated source capture
Staleness state
Dated mixed-source configurations
Question availability
Upstream tasks and terms; aggregate scores only
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Cloudflare Decision Index 0.2.1: API-Bank measure?
Published Cloudflare and community decision-model configurations.
Which model leads the published Cloudflare Decision Index 0.2.1: API-Bank snapshot?
Clef-flash currently leads the published Cloudflare Decision Index 0.2.1: API-Bank snapshot with 93.1% accuracy. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Cloudflare Decision Index 0.2.1: API-Bank?
The October 1, 2026 capture snapshot contains 64 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.