# BOOLQ (Circa) Benchmark Scores & Performance

> BOOLQ (Circa) has published results in 2 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-circa-boolq

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | Louis et al. |
| Source Type | Research system |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://aclanthology.org/2020.emnlp-main.601/) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: BOOLQ (Circa)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-circa-boolq) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-circa-boolq&format=csv)

### Relaxed labels: matched and unmatched tests

Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.

Each row retains its own training configuration. Unmatched values summarize the ten leave-one-situation-out tests.

“I’d rather just go to bed”: Understanding Indirect Answers — Annie Louis, Dan Roth, and Filip Radlinski. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2020 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2020.emnlp-main.601/) · [Full Circa results](/benchmarks/circa)

| System | Matched development accuracy (%) | Matched test accuracy (%) | Yes F-score (%) | No F-score (%) | Conditional yes F-score (%) | Middle F-score (%) | Unmatched mean accuracy (%) | Unmatched std. | Unmatched min. (%) | Unmatched max. (%) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| BOOLQ | 64.2 | 62.7 | 71.1 | 59.6 | 0.0 | 0.0 | 63.3 | 2.7 | 58.3 | 66.5 |

### Strict labels: matched and unmatched tests

Circa reports four-label relaxed and six-label strict classification separately. Matched tests use random splits; unmatched tests leave out a conversational situation. Test accuracies, class F-scores, and unmatched mean, standard deviation, minimum, and maximum are percentages. Majority-class results are baselines rather than learned model profiles.

Each row retains its own training configuration. Unmatched values summarize the ten leave-one-situation-out tests.

“I’d rather just go to bed”: Understanding Indirect Answers — Annie Louis, Dan Roth, and Filip Radlinski. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2020 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2020.emnlp-main.601/) · [Full Circa results](/benchmarks/circa)

| System | Matched development accuracy (%) | Matched test accuracy (%) | Yes F-score (%) | Probably yes F-score (%) | Conditional yes F-score (%) | No F-score (%) | Probably no F-score (%) | Middle F-score (%) | Unmatched mean accuracy (%) | Unmatched std. | Unmatched min. (%) | Unmatched max. (%) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| BOOLQ | 59.4 | 59.2 | 70.4 | 0.0 | 0.0 | 57.0 | 0.0 | 0.0 | 58.9 | 3.0 | 53.8 | 63.7 |

## Other Louis et al. Models

- [BERT-BOOLQ-YN (Circa)](/models/native-circa-bert-boolq-yn) - Score: not computed
- [BERT-DIS-YN (Circa)](/models/native-circa-bert-dis-yn) - Score: not computed
- [BERT-MNLI-YN (Circa)](/models/native-circa-bert-mnli-yn) - Score: not computed
- [BERT-YN (Answer) (Circa)](/models/native-circa-bert-yn-answer) - Score: not computed
- [BERT-YN (Circa)](/models/native-circa-bert-yn) - Score: not computed
- [BERT-YN (Question) (Circa)](/models/native-circa-bert-yn-question) - Score: not computed
- [MNLI (Circa)](/models/native-circa-mnli) - Score: not computed
