Circa
We show this table for reference; we do not rank on it.
Circa pairs polar questions with indirect answers and labels how each answer should be interpreted. It tests conversational decisions where the response does not simply say yes or no.
Original benchmark results
Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.
2 source tables · 8 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)
Relaxed labels: matched and unmatched tests
Each row retains its own training configuration. Unmatched values summarize the ten leave-one-situation-out tests.
“I’d rather just go to bed”: Understanding Indirect Answers — Annie Louis, Dan Roth, and Filip Radlinski. CC BY 4.0. Copyright 2020 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| System | Matched development accuracy (%) | Matched test accuracy (%) | Yes F-score (%) | No F-score (%) | Conditional yes F-score (%) | Middle F-score (%) | Unmatched mean accuracy (%) | Unmatched std. | Unmatched min. (%) | Unmatched max. (%) |
|---|---|---|---|---|---|---|---|---|---|---|
| Baselines (no finetuning) | ||||||||||
| Majority class | 50.2 | 49.3 | 66.0 | 0.0 | 0.0 | 0.0 | 50.4 | 4.3 | 43.6 | 56.8 |
| MNLI | 28.4 | 28.9 | 34.4 | 52.8 | 0.0 | 6.9 | 28.1 | 2.8 | 24.2 | 34.1 |
| BOOLQ | 64.2 | 62.7 | 71.1 | 59.6 | 0.0 | 0.0 | 63.3 | 2.7 | 58.3 | 66.5 |
| BERT finetuned on Question or on Answer | ||||||||||
| BERT-YN (Question only) | 56.4 | 56.0 | 63.1 | 54.1 | 9.1 | 1.0 | 53.3 | 2.9 | 48.0 | 58.4 |
| BERT-YN (Answer only) | 83.0 | 81.7 | 83.9 | 80.3 | 88.9 | 18.6 | 80.1 | 5.8 | 71.4 | 87.8 |
| BERT finetuned on Question + Answer | ||||||||||
| BERT-YN | 88.4 | 87.8 | 89.8 | 87.9 | 89.9 | 28.2 | 85.5 | 3.9 | 79.0 | 90.2 |
| BERT-MNLI-YN | 89.6 | 88.2 | 90.4 | 88.5 | 89.3 | 29.4 | 87.1 | 3.0 | 81.9 | 90.3 |
| BERT-DIS-YN | 88.0 | 87.4 | 89.4 | 87.4 | 90.0 | 35.2 | 85.5 | 3.5 | 78.9 | 89.4 |
| BERT-BOOLQ-YN | 87.7 | 87.1 | 89.0 | 86.9 | 89.6 | 30.9 | 85.3 | 3.7 | 78.8 | 89.4 |
Strict labels: matched and unmatched tests
Each row retains its own training configuration. Unmatched values summarize the ten leave-one-situation-out tests.
“I’d rather just go to bed”: Understanding Indirect Answers — Annie Louis, Dan Roth, and Filip Radlinski. CC BY 4.0. Copyright 2020 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| System | Matched development accuracy (%) | Matched test accuracy (%) | Yes F-score (%) | Probably yes F-score (%) | Conditional yes F-score (%) | No F-score (%) | Probably no F-score (%) | Middle F-score (%) | Unmatched mean accuracy (%) | Unmatched std. | Unmatched min. (%) | Unmatched max. (%) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Baselines (no finetuning) | ||||||||||||
| Majority class | 47.5 | 47.0 | 63.9 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 46.9 | 3.9 | 40.0 | 52.3 |
| MNLI | 26.3 | 27.4 | 36.6 | 0.0 | 0.0 | 53.0 | 0.0 | 4.9 | 26.4 | 3.2 | 21.7 | 32.7 |
| BOOLQ | 59.4 | 59.2 | 70.4 | 0.0 | 0.0 | 57.0 | 0.0 | 0.0 | 58.9 | 3.0 | 53.8 | 63.7 |
| BERT finetuned on Question or Answer | ||||||||||||
| BERT-YN (Question) | 53.7 | 52.8 | 62.3 | 3.2 | 19.7 | 51.1 | 0.0 | 4.7 | 49.4 | 4.0 | 41.9 | 56.7 |
| BERT-YN (Answer) | 77.3 | 77.8 | 82.5 | 49.5 | 90.2 | 77.3 | 16.2 | 26.9 | 75.8 | 5.8 | 65.4 | 82.8 |
| BERT finetuned on Question + Answer | ||||||||||||
| BERT-YN | 83.6 | 84.0 | 88.7 | 49.9 | 90.2 | 85.4 | 18.6 | 42.6 | 81.2 | 4.6 | 71.8 | 85.6 |
| BERT-MNLI-YN | 85.0 | 84.8 | 89.8 | 51.8 | 89.8 | 86.6 | 18.0 | 41.3 | 82.8 | 4.0 | 74.4 | 86.7 |
| BERT-DIS-YN | 83.8 | 83.3 | 87.9 | 50.2 | 90.5 | 84.1 | 21.2 | 50.8 | 81.5 | 4.5 | 73.1 | 86.3 |
| BERT-BOOLQ-YN | 83.1 | 83.4 | 88.2 | 51.2 | 89.1 | 84.5 | 22.1 | 43.7 | 81.1 | 4.3 | 73.3 | 85.8 |
Original benchmark results and configurations
Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.
Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.
Snapshot
About Circa
Year
2020
Tasks
Interpreting indirect answers to yes/no questions
Format
Decision classification
Difficulty
Depends on the evaluated split and label mapping
Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.
Freshness and provenance
Version
Circa 2020
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
Which Circa results are included?
This page includes 2 numeric result and configuration tables from Circa’s original paper and available benchmark-owner updates, covering 8 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.
Can I compare these Circa scores with the Perplexity panel?
Compare Circa scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.
Where can I download the Circa results?
The results download on this page provides every imported Circa table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.