Belebele (Perplexity panel)
We show this table for reference; we do not rank on it.
Passage comprehension on the 500-row sample in Perplexity's September 2026 decision panel. This is a sampled, provider-reported accuracy result with its own evaluation setup.
Provider-reported panel accuracy on Belebele (Perplexity panel) — October 1, 2026
We mirror the published provider-reported panel accuracy view for Belebele (Perplexity panel). Jev 1.13.0 leads the public snapshot at 95.00%, followed by Perplexity Decider v1 27B (94.00%) and Qwen3.8-27B (93.20%). We do not use these results to rank models overall.
Jev 1.13.0
TypeSafe AI
Perplexity panel; TypeSafe jev-1.13.0
Perplexity Decider v1 27B
Perplexity
Measured through the Perplexity API
Qwen3.8-27B
Alibaba
Perplexity panel; inference settings not specified in the model card
3 modelsDecision ModelsArchivedDisplay onlyUpdated October 1, 2026
Provider-reported panel accuracy table (3 models)
ScoreRead this as a fixed-panel result
Perplexity published this 500-row result on October 1, 2026, for its September decision panel. The three columns identify TypeSafe Jev 1.13.0, Qwen3.8-27B, and Perplexity Decider v1 27B. Decider was measured through the Perplexity API.
The sample, label conversion, and serving setup define this result. Exact prompts, item identities, baseline inference settings, and uncertainty are not specified in the model card. This panel does not replace full-dataset results.
Perplexity reports these results. We have not independently rerun the panel. Every value stays display only and outside overall and category scoring.
Snapshot
The published Belebele (Perplexity panel) snapshot places Jev 1.13.0 first at 95.00%. The third row is 1.80 points behind. The broader top-10 range is 1.80 points, so many of the published results sit in a relatively narrow band.
3 models have been evaluated on Belebele (Perplexity panel). The benchmark falls in the Decision Models category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Belebele (Perplexity panel) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Belebele (Perplexity panel)
Year
2026
Tasks
Passage comprehension
Format
Accuracy on the Perplexity fixed panel
Difficulty
Depends on the sampled tasks and label mapping
The launch chart reports 500 rows for this task. The pinned model card preserves three comparison values and states that Decider was measured through the Perplexity API. Exact prompts, item identities, label mappings, baseline inference settings, and uncertainty are not specified in the model card. Keep this result separate from full-dataset evaluations and other harnesses. These results are display only and excluded from overall and category scoring. JevBench public-hard accuracy is separate from the full JevBench composite, and TruthfulQA binary is a separate label conversion.
Freshness and provenance
Version
September 2026 Perplexity decision panel
Refresh cadence
Pinned provider report
Staleness state
Archived
Question availability
Task families public; sampled item identities and prompts not specified in the model card
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
Which systems are compared in the Belebele panel?
Perplexity compares TypeSafe Jev 1.13.0, Qwen3.8-27B, and pplx-decider-v1-27b on 500 rows. The chart covers September 2026 and was published October 1. Decider was measured through the Perplexity API. The model card does not specify the baselines’ inference settings, so the scores describe this particular comparison.
Are these independently verified benchmark results?
Perplexity published the percentages in its model card. We checked the values against that pinned source, but we have not independently rerun the evaluation. Provider-reported means the values match the provider’s report; it does not mean an independent evaluator reproduced the scores or established their uncertainty.
Can these scores replace full benchmark results?
Keep the fixed-panel scores separate. Sample sizes, label conversions, prompts, and serving settings can change accuracy, and the model card does not specify enough detail to match another harness. The panel measures correctness on these samples. It does not measure probability calibration or enter overall model rankings.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.