ContractNLI
We show this table for reference; we do not rank on it.
ContractNLI asks whether a contract entails, contradicts, or does not mention a hypothesis. The original task also requires evidence spans that support the classification.
Original benchmark results
ContractNLI evaluates both document-level three-class inference and evidence retrieval. Accuracy, F1, mAP, and P@R80 are 0–1 ratios. Means and standard deviations occupy separate columns. Oracle evidence experiments use a distinct binary subset. We preserve the paper’s differing BERT-large text-hypothesis accuracy deviations rather than choosing one.
7 source tables · 19 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)
Backbones and pretraining
All measurements are 0–1 ratios, with standard deviations alongside means. “None” means no additional domain pretraining before ContractNLI training.
ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Backbone | Pretraining or fine-tuning | Evidence mAP mean | mAP std. | Evidence P@R80 mean | P@R80 std. | NLI accuracy mean | Accuracy std. | Contradiction F1 mean | Contradiction F1 std. | Entailment F1 mean | Entailment F1 std. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BERT base | None | .885 | .025 | .663 | .093 | .838 | .020 | .287 | .022 | .765 | .035 |
| BERT large | None | .922 | .006 | .793 | .018 | .875 | .006 | .357 | .039 | .834 | .002 |
| DeBERTa v2 xlarge | None | .933 | .002 | .859 | .008 | .885 | .001 | .360 | .027 | .855 | .002 |
| BERT base | Pretrained from scratch using a case law corpus Zheng et al. (2021) | .870 | .015 | .578 | .052 | .831 | .032 | .289 | .026 | .783 | .040 |
| BERT base | Fine-tuned on case law and contract corpora Chalkidis et al. (2020) | .925 | .004 | .811 | .002 | .794 | .008 | .272 | .008 | .746 | .018 |
| DeBERTa v2 xlarge | Fine-tuned on span identification Hendrycks et al. (2021) | .936 | .002 | .860 | .003 | .892 | .001 | .405 | .016 | .859 | .005 |
| BERT base | Fine-tuned on NDAs | .892 | .002 | .690 | .014 | .864 | .004 | .326 | .014 | .820 | .010 |
| BERT large | Fine-tuned on NDAs | .922 | .003 | .837 | .008 | .875 | .000 | .389 | .009 | .839 | .003 |
Original baselines and proposed models
Blank deviations are unreported. Means and standard deviations are retained without conversion to percentage points.
ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| System | Evidence mAP mean | mAP std. | Evidence P@R80 mean | P@R80 std. | NLI accuracy mean | Accuracy std. | Contradiction F1 mean | Contradiction F1 std. | Entailment F1 mean | Entailment F1 std. |
|---|---|---|---|---|---|---|---|---|---|---|
| Majority vote | — | Unreported | — | Unreported | .674 | Unreported | .083 | Unreported | .428 | Unreported |
| Doc TF-IDF+SVM | — | Unreported | — | Unreported | .733 | Unreported | .197 | Unreported | .641 | Unreported |
| Random | .024 | Unreported | .000 | Unreported | — | Unreported | — | Unreported | — | Unreported |
| Span TF-IDF+Cosine | .381 | Unreported | .057 | Unreported | — | Unreported | — | Unreported | — | Unreported |
| Span TF-IDF+SVM | .836 | Unreported | .322 | Unreported | — | Unreported | — | Unreported | — | Unreported |
| SQuAD (BERT base ) | .825 | .004 | .574 | .004 | — | Unreported | — | Unreported | — | Unreported |
| SQuAD (BERT large ) | .869 | .005 | .661 | .043 | — | Unreported | — | Unreported | — | Unreported |
| Ours (BERT base ) | .885 | .025 | .663 | .093 | .838 | .020 | .287 | .022 | .765 | .035 |
| Ours (BERT large ) | .922 | .006 | .793 | .018 | .875 | .006 | .357 | .039 | .834 | .002 |
Symbol versus text hypotheses
Blank deviations are unreported. Means and standard deviations are retained without conversion to percentage points.
ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| System | Evidence mAP mean | mAP std. | Evidence P@R80 mean | P@R80 std. | NLI accuracy mean | Accuracy std. | Contradiction F1 mean | Contradiction F1 std. | Entailment F1 mean | Entailment F1 std. |
|---|---|---|---|---|---|---|---|---|---|---|
| Symbol (BERT base ) | .857 | .044 | .574 | .136 | .830 | .014 | .294 | .075 | .751 | .027 |
| Symbol (BERT large ) | .894 | .020 | .703 | .092 | .849 | .006 | .303 | .058 | .794 | .026 |
| Text (BERT base ) | .885 | .025 | .663 | .093 | .838 | .020 | .287 | .022 | .765 | .035 |
| Text (BERT large ) | .922 | .006 | .793 | .018 | .875 | .016 | .357 | .039 | .834 | .002 |
Binary subset with predicted or oracle evidence
Binary subset: not comparable with the main three-class document test. Oracle uses gold evidence and is not a deployable model result.
ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| System | NLI accuracy mean | Accuracy std. | Contradiction F1 mean | Contradiction F1 std. | Entailment F1 mean | Entailment F1 std. |
|---|---|---|---|---|---|---|
| Majority vote | .814 | Unreported | .239 | Unreported | .645 | Unreported |
| Span NLI (BERT base ) | .883 | .006 | .490 | .007 | .795 | .005 |
| Span NLI (BERT large ) | .899 | .004 | .492 | .065 | .820 | .012 |
| Oracle NLI (BERT base ) | .918 | .005 | .657 | .062 | .816 | .006 |
| Oracle NLI (BERT large ) | .908 | .011 | .620 | .082 | .806 | .015 |
Negation-condition ablation
Paper ablation; numeric cells remain in their source units and cannot be pooled with the main test.
ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Condition | Majority-label NLI accuracy | Minority-label accuracy | Weighted accuracy | Minority-label share (%) |
|---|---|---|---|---|
| w/o (local) | .91 | .77 | .84 | 21 |
| w/ (local) | .92 | .40 | .66 | 7 |
| w/o (non-local) | .98 | .72 | .85 | 19 |
| w/ (non-local) | .90 | .00 | .45 | 6 |
Continuous and discontinuous evidence
Paper ablation; numeric cells remain in their source units and cannot be pooled with the main test.
ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Evidence shape | Context n | Mean spans | Spans before one evidence span | Spans before all evidence | Evidence mAP |
|---|---|---|---|---|---|
| Continuous | 128 | 2.64 | 1.09 | 3.82 | 0.91 |
| Discontinuous | 128 | 2.34 | 1.04 | 3.84 | 0.94 |
| Continuous | 64 | 2.64 | 1.16 | 4.33 | 0.89 |
| Discontinuous | 64 | 2.34 | 1.01 | 4.85 | 0.94 |
Reference-condition ablation
Paper ablation; numeric cells remain in their source units and cannot be pooled with the main test.
ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Condition | Majority-label NLI accuracy | Minority-label accuracy | Weighted accuracy | Minority-label share (%) |
|---|---|---|---|---|
| w/o Reference | .91 | .88 | .89 | 26 |
| w/ Reference | .93 | — | — | 0 |
Original benchmark results and configurations
ContractNLI evaluates both document-level three-class inference and evidence retrieval. Accuracy, F1, mAP, and P@R80 are 0–1 ratios. Means and standard deviations occupy separate columns. Oracle evidence experiments use a distinct binary subset. We preserve the paper’s differing BERT-large text-hypothesis accuracy deviations rather than choosing one.
Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.
Snapshot
About ContractNLI
Year
2021
Tasks
Contract entailment and evidence identification
Format
Decision classification
Difficulty
Depends on the evaluated split and label mapping
ContractNLI evaluates both document-level three-class inference and evidence retrieval. Accuracy, F1, mAP, and P@R80 are 0–1 ratios. Means and standard deviations occupy separate columns. Oracle evidence experiments use a distinct binary subset. We preserve the paper’s differing BERT-large text-hypothesis accuracy deviations rather than choosing one.
Freshness and provenance
Version
ContractNLI 2021
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
Which ContractNLI results are included?
This page includes 7 numeric result and configuration tables from ContractNLI’s original paper and available benchmark-owner updates, covering 19 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.
Can I compare these ContractNLI scores with the Perplexity panel?
Compare ContractNLI scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.
Where can I download the ContractNLI results?
The results download on this page provides every imported ContractNLI table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.