# ContractNLI

> ContractNLI asks whether a contract entails, contradicts, or does not mention a hypothesis. The original task also requires evidence spans that support the classification.

Canonical page: https://benchlm.ai/benchmarks/contractnli

- Category: [Decision Models](/decision-models)
- Last updated: October 1, 2026 source review

## About ContractNLI

- Year: 2021
- Tasks: Contract entailment and evidence identification
- Format: Decision classification
- Difficulty: Depends on the evaluated split and label mapping
- Paper: [ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts](https://stanfordnlp.github.io/contract-nli/)

ContractNLI evaluates both document-level three-class inference and evidence retrieval. Accuracy, F1, mAP, and P@R80 are 0–1 ratios. Means and standard deviations occupy separate columns. Oracle evidence experiments use a distinct binary subset. We preserve the paper’s differing BERT-large text-hypothesis accuracy deviations rather than choosing one.

ContractNLI is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Original benchmark results

ContractNLI evaluates both document-level three-class inference and evidence retrieval. Accuracy, F1, mAP, and P@R80 are 0–1 ratios. Means and standard deviations occupy separate columns. Oracle evidence experiments use a distinct binary subset. We preserve the paper’s differing BERT-large text-hypothesis accuracy deviations rather than choosing one.

[Download full results (JSON)](/api/data/decision-benchmarks?benchmark=contractNli)

### Backbones and pretraining

All measurements are 0–1 ratios, with standard deviations alongside means. “None” means no additional domain pretraining before ContractNLI training.

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2021.findings-emnlp.164/)

| Backbone | Pretraining or fine-tuning | Evidence mAP mean | mAP std. | Evidence P@R80 mean | P@R80 std. | NLI accuracy mean | Accuracy std. | Contradiction F1 mean | Contradiction F1 std. | Entailment F1 mean | Entailment F1 std. |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [BERT base](/models/native-contractnli-bert-base-none) | None | .885 | .025 | .663 | .093 | .838 | .020 | .287 | .022 | .765 | .035 |
| [BERT large](/models/native-contractnli-bert-large-none) | None | .922 | .006 | .793 | .018 | .875 | .006 | .357 | .039 | .834 | .002 |
| [DeBERTa v2 xlarge](/models/native-contractnli-deberta-v2-xlarge-none) | None | .933 | .002 | .859 | .008 | .885 | .001 | .360 | .027 | .855 | .002 |
| [BERT base](/models/native-contractnli-bert-base-pretrained-from-scratch-using-a-case-law-corpus-zheng-et-al-2021) | Pretrained from scratch using a case law corpus Zheng et al. (2021) | .870 | .015 | .578 | .052 | .831 | .032 | .289 | .026 | .783 | .040 |
| [BERT base](/models/native-contractnli-bert-base-fine-tuned-on-case-law-and-contract-corpora-chalkidis-et-al-2020) | Fine-tuned on case law and contract corpora Chalkidis et al. (2020) | .925 | .004 | .811 | .002 | .794 | .008 | .272 | .008 | .746 | .018 |
| [DeBERTa v2 xlarge](/models/native-contractnli-deberta-v2-xlarge-fine-tuned-on-span-identification-hendrycks-et-al-2021) | Fine-tuned on span identification Hendrycks et al. (2021) | .936 | .002 | .860 | .003 | .892 | .001 | .405 | .016 | .859 | .005 |
| [BERT base](/models/native-contractnli-bert-base-fine-tuned-on-ndas) | Fine-tuned on NDAs | .892 | .002 | .690 | .014 | .864 | .004 | .326 | .014 | .820 | .010 |
| [BERT large](/models/native-contractnli-bert-large-fine-tuned-on-ndas) | Fine-tuned on NDAs | .922 | .003 | .837 | .008 | .875 | .000 | .389 | .009 | .839 | .003 |

### Original baselines and proposed models

Blank deviations are unreported. Means and standard deviations are retained without conversion to percentage points.

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2021.findings-emnlp.164/)

| System | Evidence mAP mean | mAP std. | Evidence P@R80 mean | P@R80 std. | NLI accuracy mean | Accuracy std. | Contradiction F1 mean | Contradiction F1 std. | Entailment F1 mean | Entailment F1 std. |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Majority vote | — |  | — |  | .674 |  | .083 |  | .428 |  |
| [Doc TF-IDF+SVM](/models/native-contractnli-doc-tf-idf-svm) | — |  | — |  | .733 |  | .197 |  | .641 |  |
| Random | .024 |  | .000 |  | — |  | — |  | — |  |
| [Span TF-IDF+Cosine](/models/native-contractnli-span-tf-idf-cosine) | .381 |  | .057 |  | — |  | — |  | — |  |
| [Span TF-IDF+SVM](/models/native-contractnli-span-tf-idf-svm) | .836 |  | .322 |  | — |  | — |  | — |  |
| [SQuAD (BERT base )](/models/native-contractnli-squad-bert-base) | .825 | .004 | .574 | .004 | — |  | — |  | — |  |
| [SQuAD (BERT large )](/models/native-contractnli-squad-bert-large) | .869 | .005 | .661 | .043 | — |  | — |  | — |  |
| [Ours (BERT base )](/models/native-contractnli-bert-base-none) | .885 | .025 | .663 | .093 | .838 | .020 | .287 | .022 | .765 | .035 |
| [Ours (BERT large )](/models/native-contractnli-bert-large-none) | .922 | .006 | .793 | .018 | .875 | .006 | .357 | .039 | .834 | .002 |

### Symbol versus text hypotheses

Blank deviations are unreported. Means and standard deviations are retained without conversion to percentage points.

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2021.findings-emnlp.164/)

| System | Evidence mAP mean | mAP std. | Evidence P@R80 mean | P@R80 std. | NLI accuracy mean | Accuracy std. | Contradiction F1 mean | Contradiction F1 std. | Entailment F1 mean | Entailment F1 std. |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Symbol (BERT base )](/models/native-contractnli-symbol-bert-base) | .857 | .044 | .574 | .136 | .830 | .014 | .294 | .075 | .751 | .027 |
| [Symbol (BERT large )](/models/native-contractnli-symbol-bert-large) | .894 | .020 | .703 | .092 | .849 | .006 | .303 | .058 | .794 | .026 |
| [Text (BERT base )](/models/native-contractnli-bert-base-none) | .885 | .025 | .663 | .093 | .838 | .020 | .287 | .022 | .765 | .035 |
| [Text (BERT large )](/models/native-contractnli-bert-large-none) | .922 | .006 | .793 | .018 | .875 | .016 | .357 | .039 | .834 | .002 |

### Binary subset with predicted or oracle evidence

Binary subset: not comparable with the main three-class document test. Oracle uses gold evidence and is not a deployable model result.

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2021.findings-emnlp.164/)

| System | NLI accuracy mean | Accuracy std. | Contradiction F1 mean | Contradiction F1 std. | Entailment F1 mean | Entailment F1 std. |
| --- | --- | --- | --- | --- | --- | --- |
| Majority vote | .814 |  | .239 |  | .645 |  |
| [Span NLI (BERT base )](/models/native-contractnli-span-nli-bert-base-binary-subset) | .883 | .006 | .490 | .007 | .795 | .005 |
| [Span NLI (BERT large )](/models/native-contractnli-span-nli-bert-large-binary-subset) | .899 | .004 | .492 | .065 | .820 | .012 |
| [Oracle NLI (BERT base )](/models/native-contractnli-oracle-nli-bert-base-binary-subset) | .918 | .005 | .657 | .062 | .816 | .006 |
| [Oracle NLI (BERT large )](/models/native-contractnli-oracle-nli-bert-large-binary-subset) | .908 | .011 | .620 | .082 | .806 | .015 |

### Negation-condition ablation

Paper ablation; numeric cells remain in their source units and cannot be pooled with the main test.

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2021.findings-emnlp.164/)

| Condition | Majority-label NLI accuracy | Minority-label accuracy | Weighted accuracy | Minority-label share (%) |
| --- | --- | --- | --- | --- |
| w/o (local) | .91 | .77 | .84 | 21 |
| w/ (local) | .92 | .40 | .66 | 7 |
| w/o (non-local) | .98 | .72 | .85 | 19 |
| w/ (non-local) | .90 | .00 | .45 | 6 |

### Continuous and discontinuous evidence

Paper ablation; numeric cells remain in their source units and cannot be pooled with the main test.

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2021.findings-emnlp.164/)

| Evidence shape | Context n | Mean spans | Spans before one evidence span | Spans before all evidence | Evidence mAP |
| --- | --- | --- | --- | --- | --- |
| Continuous | 128 | 2.64 | 1.09 | 3.82 | 0.91 |
| Discontinuous | 128 | 2.34 | 1.04 | 3.84 | 0.94 |
| Continuous | 64 | 2.64 | 1.16 | 4.33 | 0.89 |
| Discontinuous | 64 | 2.34 | 1.01 | 4.85 | 0.94 |

### Reference-condition ablation

Paper ablation; numeric cells remain in their source units and cannot be pooled with the main test.

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2021.findings-emnlp.164/)

| Condition | Majority-label NLI accuracy | Minority-label accuracy | Weighted accuracy | Minority-label share (%) |
| --- | --- | --- | --- | --- |
| w/o Reference | .91 | .88 | .89 | 26 |
| w/ Reference | .93 | — | — | 0 |

## FAQ

### Which ContractNLI results are included?

This page includes 7 numeric result and configuration tables from ContractNLI’s original paper and available benchmark-owner updates, covering 19 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

### Can I compare these ContractNLI scores with the Perplexity panel?

Compare ContractNLI scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

### Where can I download the ContractNLI results?

The results download on this page provides every imported ContractNLI table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
