# Perplexity Decision Panel

> Perplexity reports accuracy across 11 decision tasks on a fixed 7,210-row panel. The overall percentage combines unequal task sample sizes, so a larger task contributes more than a smaller task.

Canonical page: https://benchlm.ai/benchmarks/perplexity-decision-panel

- Category: [Decision Models](/decision-models)
- Last updated: October 1, 2026

## About Perplexity Decision Panel

- Year: 2026
- Tasks: 11 decision task families
- Format: Accuracy on the Perplexity fixed panel
- Difficulty: Depends on the sampled tasks and label mapping
- Paper: [Perplexity Decider v1 27B model card](https://huggingface.co/perplexity-ai/pplx-decider-v1-27b/blob/9ce1abcf1f00209405376b5bcc81225c8f8cf514/README.md)

The October 1 model card publishes three comparison columns: TypeSafe Jev 1.13.0, Qwen3.8-27B, and pplx-decider-v1-27b. Perplexity says its Decider results were measured through its API. The launch chart gives the sample sizes; the pinned card confirms every percentage. The published overall is retained without rescoring. These results are display only and excluded from overall and category scoring. JevBench public-hard accuracy is separate from the full JevBench composite, and TruthfulQA binary is a separate label conversion.

Perplexity Decision Panel is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Perplexity Decider v1 27B](/models/pplx-decider-v1-27b) | Measured through the Perplexity API | Perplexity | 85.71% |
| 2 | [Jev 1.13.0](/models/jev-1-13-0) | Perplexity panel; TypeSafe jev-1.13.0 | TypeSafe AI | 84.51% |
| 3 | [Qwen3.8-27B](/models/qwen3-8-27b) | Perplexity panel; inference settings not specified in the model card | Alibaba | 74.76% |

## FAQ

### Which systems are compared in the Perplexity decision panel?

Perplexity compares TypeSafe Jev 1.13.0, Qwen3.8-27B, and pplx-decider-v1-27b on 7,210 rows. The chart covers September 2026 and was published October 1. Decider was measured through the Perplexity API. The model card does not specify the baselines’ inference settings, so the scores describe this particular comparison.

### Are these independently verified benchmark results?

Perplexity published the percentages in its model card. We checked the values against that pinned source, but we have not independently rerun the evaluation. Provider-reported means the values match the provider’s report; it does not mean an independent evaluator reproduced the scores or established their uncertainty.

### Can these scores replace full benchmark results?

Keep the fixed-panel scores separate. Sample sizes, label conversions, prompts, and serving settings can change accuracy, and the model card does not specify enough detail to match another harness. The panel measures correctness on these samples. It does not measure probability calibration or enter overall model rankings.
