# Circa

> Circa pairs polar questions with indirect answers and labels how each answer should be interpreted. It tests conversational decisions where the response does not simply say yes or no.

Canonical page: https://benchlm.ai/benchmarks/circa

- Category: [Decision Models](/decision-models)
- Last updated: October 1, 2026 source review

## About Circa

- Year: 2020
- Tasks: Interpreting indirect answers to yes/no questions
- Format: Decision classification
- Difficulty: Depends on the evaluated split and label mapping
- Paper: [“I’d rather just go to bed”: Understanding Indirect Answers](https://github.com/google-research-datasets/circa)

Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.

Circa is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Original benchmark results

Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.

[Download full results (JSON)](/api/data/decision-benchmarks?benchmark=circa)

### Relaxed labels: matched and unmatched tests

Each row retains its own training configuration. Unmatched values summarize the ten leave-one-situation-out tests.

“I’d rather just go to bed”: Understanding Indirect Answers — Annie Louis, Dan Roth, and Filip Radlinski. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2020 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2020.emnlp-main.601/)

| System | Matched development accuracy (%) | Matched test accuracy (%) | Yes F-score (%) | No F-score (%) | Conditional yes F-score (%) | Middle F-score (%) | Unmatched mean accuracy (%) | Unmatched std. | Unmatched min. (%) | Unmatched max. (%) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Baselines (no finetuning) |  |  |  |  |  |  |  |  |  |  |
| Majority class | 50.2 | 49.3 | 66.0 | 0.0 | 0.0 | 0.0 | 50.4 | 4.3 | 43.6 | 56.8 |
| [MNLI](/models/native-circa-mnli) | 28.4 | 28.9 | 34.4 | 52.8 | 0.0 | 6.9 | 28.1 | 2.8 | 24.2 | 34.1 |
| [BOOLQ](/models/native-circa-boolq) | 64.2 | 62.7 | 71.1 | 59.6 | 0.0 | 0.0 | 63.3 | 2.7 | 58.3 | 66.5 |
| BERT finetuned on Question or on Answer |  |  |  |  |  |  |  |  |  |  |
| [BERT-YN (Question only)](/models/native-circa-bert-yn-question) | 56.4 | 56.0 | 63.1 | 54.1 | 9.1 | 1.0 | 53.3 | 2.9 | 48.0 | 58.4 |
| [BERT-YN (Answer only)](/models/native-circa-bert-yn-answer) | 83.0 | 81.7 | 83.9 | 80.3 | 88.9 | 18.6 | 80.1 | 5.8 | 71.4 | 87.8 |
| BERT finetuned on Question + Answer |  |  |  |  |  |  |  |  |  |  |
| [BERT-YN](/models/native-circa-bert-yn) | 88.4 | 87.8 | 89.8 | 87.9 | 89.9 | 28.2 | 85.5 | 3.9 | 79.0 | 90.2 |
| [BERT-MNLI-YN](/models/native-circa-bert-mnli-yn) | 89.6 | 88.2 | 90.4 | 88.5 | 89.3 | 29.4 | 87.1 | 3.0 | 81.9 | 90.3 |
| [BERT-DIS-YN](/models/native-circa-bert-dis-yn) | 88.0 | 87.4 | 89.4 | 87.4 | 90.0 | 35.2 | 85.5 | 3.5 | 78.9 | 89.4 |
| [BERT-BOOLQ-YN](/models/native-circa-bert-boolq-yn) | 87.7 | 87.1 | 89.0 | 86.9 | 89.6 | 30.9 | 85.3 | 3.7 | 78.8 | 89.4 |

### Strict labels: matched and unmatched tests

Each row retains its own training configuration. Unmatched values summarize the ten leave-one-situation-out tests.

“I’d rather just go to bed”: Understanding Indirect Answers — Annie Louis, Dan Roth, and Filip Radlinski. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2020 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2020.emnlp-main.601/)

| System | Matched development accuracy (%) | Matched test accuracy (%) | Yes F-score (%) | Probably yes F-score (%) | Conditional yes F-score (%) | No F-score (%) | Probably no F-score (%) | Middle F-score (%) | Unmatched mean accuracy (%) | Unmatched std. | Unmatched min. (%) | Unmatched max. (%) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Baselines (no finetuning) |  |  |  |  |  |  |  |  |  |  |  |  |
| Majority class | 47.5 | 47.0 | 63.9 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 46.9 | 3.9 | 40.0 | 52.3 |
| [MNLI](/models/native-circa-mnli) | 26.3 | 27.4 | 36.6 | 0.0 | 0.0 | 53.0 | 0.0 | 4.9 | 26.4 | 3.2 | 21.7 | 32.7 |
| [BOOLQ](/models/native-circa-boolq) | 59.4 | 59.2 | 70.4 | 0.0 | 0.0 | 57.0 | 0.0 | 0.0 | 58.9 | 3.0 | 53.8 | 63.7 |
| BERT finetuned on Question or Answer |  |  |  |  |  |  |  |  |  |  |  |  |
| [BERT-YN (Question)](/models/native-circa-bert-yn-question) | 53.7 | 52.8 | 62.3 | 3.2 | 19.7 | 51.1 | 0.0 | 4.7 | 49.4 | 4.0 | 41.9 | 56.7 |
| [BERT-YN (Answer)](/models/native-circa-bert-yn-answer) | 77.3 | 77.8 | 82.5 | 49.5 | 90.2 | 77.3 | 16.2 | 26.9 | 75.8 | 5.8 | 65.4 | 82.8 |
| BERT finetuned on Question + Answer |  |  |  |  |  |  |  |  |  |  |  |  |
| [BERT-YN](/models/native-circa-bert-yn) | 83.6 | 84.0 | 88.7 | 49.9 | 90.2 | 85.4 | 18.6 | 42.6 | 81.2 | 4.6 | 71.8 | 85.6 |
| [BERT-MNLI-YN](/models/native-circa-bert-mnli-yn) | 85.0 | 84.8 | 89.8 | 51.8 | 89.8 | 86.6 | 18.0 | 41.3 | 82.8 | 4.0 | 74.4 | 86.7 |
| [BERT-DIS-YN](/models/native-circa-bert-dis-yn) | 83.8 | 83.3 | 87.9 | 50.2 | 90.5 | 84.1 | 21.2 | 50.8 | 81.5 | 4.5 | 73.1 | 86.3 |
| [BERT-BOOLQ-YN](/models/native-circa-bert-boolq-yn) | 83.1 | 83.4 | 88.2 | 51.2 | 89.1 | 84.5 | 22.1 | 43.7 | 81.1 | 4.3 | 73.3 | 85.8 |

## FAQ

### Which Circa results are included?

This page includes 2 numeric result and configuration tables from Circa’s original paper and available benchmark-owner updates, covering 8 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

### Can I compare these Circa scores with the Perplexity panel?

Compare Circa scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

### Where can I download the Circa results?

The results download on this page provides every imported Circa table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
