# WinoGrande (Perplexity panel)

> Ambiguous-reference resolution on the 1,000-row sample in Perplexity's September 2026 decision panel. This is a sampled, provider-reported accuracy result with its own evaluation setup.

Canonical page: https://benchlm.ai/benchmarks/winogrande-perplexity-panel

- Category: [Decision Models](/decision-models)
- Last updated: October 1, 2026

## About WinoGrande (Perplexity panel)

- Year: 2026
- Tasks: Ambiguous-reference resolution
- Format: Accuracy on the Perplexity fixed panel
- Difficulty: Depends on the sampled tasks and label mapping
- Paper: [Perplexity Decider v1 27B model card](https://huggingface.co/perplexity-ai/pplx-decider-v1-27b/blob/9ce1abcf1f00209405376b5bcc81225c8f8cf514/README.md)

The launch chart reports 1,000 rows for this task. The pinned model card preserves three comparison values and states that Decider was measured through the Perplexity API. Exact prompts, item identities, label mappings, baseline inference settings, and uncertainty are not specified in the model card. Keep this result separate from full-dataset evaluations and other harnesses. These results are display only and excluded from overall and category scoring. JevBench public-hard accuracy is separate from the full JevBench composite, and TruthfulQA binary is a separate label conversion.

WinoGrande (Perplexity panel) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Jev 1.13.0](/models/jev-1-13-0) | Perplexity panel; TypeSafe jev-1.13.0 | TypeSafe AI | 90.70% |
| 2 | [Perplexity Decider v1 27B](/models/pplx-decider-v1-27b) | Measured through the Perplexity API | Perplexity | 83.30% |
| 3 | [Qwen3.8-27B](/models/qwen3-8-27b) | Perplexity panel; inference settings not specified in the model card | Alibaba | 73.10% |

## FAQ

### Which systems are compared in the WinoGrande panel?

Perplexity compares TypeSafe Jev 1.13.0, Qwen3.8-27B, and pplx-decider-v1-27b on 1,000 rows. The chart covers September 2026 and was published October 1. Decider was measured through the Perplexity API. The model card does not specify the baselines’ inference settings, so the scores describe this particular comparison.

### Are these independently verified benchmark results?

Perplexity published the percentages in its model card. We checked the values against that pinned source, but we have not independently rerun the evaluation. Provider-reported means the values match the provider’s report; it does not mean an independent evaluator reproduced the scores or established their uncertainty.

### Can these scores replace full benchmark results?

Keep the fixed-panel scores separate. Sample sizes, label conversions, prompts, and serving settings can change accuracy, and the model card does not specify enough detail to match another harness. The panel measures correctness on these samples. It does not measure probability calibration or enter overall model rankings.
