Skip to main content
BenchLM

Circa

We show this table for reference; we do not rank on it.

Circa pairs polar questions with indirect answers and labels how each answer should be interpreted. It tests conversational decisions where the response does not simply say yes or no.

Original benchmark results

Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.

2 source tables · 8 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)

Relaxed labels: matched and unmatched tests

Each row retains its own training configuration. Unmatched values summarize the ten leave-one-situation-out tests.

Published source

“I’d rather just go to bed”: Understanding Indirect Answers — Annie Louis, Dan Roth, and Filip Radlinski. CC BY 4.0. Copyright 2020 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Relaxed labels: matched and unmatched tests. Published source metrics with original precision; unreported cells are not zero.
SystemMatched development accuracy (%)Matched test accuracy (%)Yes F-score (%)No F-score (%)Conditional yes F-score (%)Middle F-score (%)Unmatched mean accuracy (%)Unmatched std.Unmatched min. (%)Unmatched max. (%)
Baselines (no finetuning)
Majority class50.249.366.00.00.00.050.44.343.656.8
MNLI28.428.934.452.80.06.928.12.824.234.1
BOOLQ64.262.771.159.60.00.063.32.758.366.5
BERT finetuned on Question or on Answer
BERT-YN (Question only)56.456.063.154.19.11.053.32.948.058.4
BERT-YN (Answer only)83.081.783.980.388.918.680.15.871.487.8
BERT finetuned on Question + Answer
BERT-YN88.487.889.887.989.928.285.53.979.090.2
BERT-MNLI-YN89.688.290.488.589.329.487.13.081.990.3
BERT-DIS-YN88.087.489.487.490.035.285.53.578.989.4
BERT-BOOLQ-YN87.787.189.086.989.630.985.33.778.889.4

Strict labels: matched and unmatched tests

Each row retains its own training configuration. Unmatched values summarize the ten leave-one-situation-out tests.

Published source

“I’d rather just go to bed”: Understanding Indirect Answers — Annie Louis, Dan Roth, and Filip Radlinski. CC BY 4.0. Copyright 2020 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Strict labels: matched and unmatched tests. Published source metrics with original precision; unreported cells are not zero.
SystemMatched development accuracy (%)Matched test accuracy (%)Yes F-score (%)Probably yes F-score (%)Conditional yes F-score (%)No F-score (%)Probably no F-score (%)Middle F-score (%)Unmatched mean accuracy (%)Unmatched std.Unmatched min. (%)Unmatched max. (%)
Baselines (no finetuning)
Majority class47.547.063.90.00.00.00.00.046.93.940.052.3
MNLI26.327.436.60.00.053.00.04.926.43.221.732.7
BOOLQ59.459.270.40.00.057.00.00.058.93.053.863.7
BERT finetuned on Question or Answer
BERT-YN (Question)53.752.862.33.219.751.10.04.749.44.041.956.7
BERT-YN (Answer)77.377.882.549.590.277.316.226.975.85.865.482.8
BERT finetuned on Question + Answer
BERT-YN83.684.088.749.990.285.418.642.681.24.671.885.6
BERT-MNLI-YN85.084.889.851.889.886.618.041.382.84.074.486.7
BERT-DIS-YN83.883.387.950.290.584.121.250.881.54.573.186.3
BERT-BOOLQ-YN83.183.488.251.289.184.522.143.781.14.373.385.8

Original benchmark results and configurations

Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.

Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.

Snapshot

2 source tables8 configurationsDisplay only

About Circa

Year

2020

Tasks

Interpreting indirect answers to yes/no questions

Format

Decision classification

Difficulty

Depends on the evaluated split and label mapping

Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.

Freshness and provenance

Version

Circa 2020

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

Which Circa results are included?

This page includes 2 numeric result and configuration tables from Circa’s original paper and available benchmark-owner updates, covering 8 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

Can I compare these Circa scores with the Perplexity panel?

Compare Circa scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

Where can I download the Circa results?

The results download on this page provides every imported Circa table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.

Last updated: October 1, 2026 source review · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.