Skip to main content
BenchLM

ContractNLI

We show this table for reference; we do not rank on it.

ContractNLI asks whether a contract entails, contradicts, or does not mention a hypothesis. The original task also requires evidence spans that support the classification.

Original benchmark results

ContractNLI evaluates both document-level three-class inference and evidence retrieval. Accuracy, F1, mAP, and P@R80 are 0–1 ratios. Means and standard deviations occupy separate columns. Oracle evidence experiments use a distinct binary subset. We preserve the paper’s differing BERT-large text-hypothesis accuracy deviations rather than choosing one.

7 source tables · 19 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)

Backbones and pretraining

All measurements are 0–1 ratios, with standard deviations alongside means. “None” means no additional domain pretraining before ContractNLI training.

Published source

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Backbones and pretraining. Published source metrics with original precision; unreported cells are not zero.
BackbonePretraining or fine-tuningEvidence mAP meanmAP std.Evidence P@R80 meanP@R80 std.NLI accuracy meanAccuracy std.Contradiction F1 meanContradiction F1 std.Entailment F1 meanEntailment F1 std.
BERT baseNone.885.025.663.093.838.020.287.022.765.035
BERT largeNone.922.006.793.018.875.006.357.039.834.002
DeBERTa v2 xlargeNone.933.002.859.008.885.001.360.027.855.002
BERT basePretrained from scratch using a case law corpus Zheng et al. (2021).870.015.578.052.831.032.289.026.783.040
BERT baseFine-tuned on case law and contract corpora Chalkidis et al. (2020).925.004.811.002.794.008.272.008.746.018
DeBERTa v2 xlargeFine-tuned on span identification Hendrycks et al. (2021).936.002.860.003.892.001.405.016.859.005
BERT baseFine-tuned on NDAs.892.002.690.014.864.004.326.014.820.010
BERT largeFine-tuned on NDAs.922.003.837.008.875.000.389.009.839.003

Original baselines and proposed models

Blank deviations are unreported. Means and standard deviations are retained without conversion to percentage points.

Published source

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Original baselines and proposed models. Published source metrics with original precision; unreported cells are not zero.
SystemEvidence mAP meanmAP std.Evidence P@R80 meanP@R80 std.NLI accuracy meanAccuracy std.Contradiction F1 meanContradiction F1 std.Entailment F1 meanEntailment F1 std.
Majority vote—Unreported—Unreported.674Unreported.083Unreported.428Unreported
Doc TF-IDF+SVM—Unreported—Unreported.733Unreported.197Unreported.641Unreported
Random.024Unreported.000Unreported—Unreported—Unreported—Unreported
Span TF-IDF+Cosine.381Unreported.057Unreported—Unreported—Unreported—Unreported
Span TF-IDF+SVM.836Unreported.322Unreported—Unreported—Unreported—Unreported
SQuAD (BERT base ).825.004.574.004—Unreported—Unreported—Unreported
SQuAD (BERT large ).869.005.661.043—Unreported—Unreported—Unreported
Ours (BERT base ).885.025.663.093.838.020.287.022.765.035
Ours (BERT large ).922.006.793.018.875.006.357.039.834.002

Symbol versus text hypotheses

Blank deviations are unreported. Means and standard deviations are retained without conversion to percentage points.

Published source

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Symbol versus text hypotheses. Published source metrics with original precision; unreported cells are not zero.
SystemEvidence mAP meanmAP std.Evidence P@R80 meanP@R80 std.NLI accuracy meanAccuracy std.Contradiction F1 meanContradiction F1 std.Entailment F1 meanEntailment F1 std.
Symbol (BERT base ).857.044.574.136.830.014.294.075.751.027
Symbol (BERT large ).894.020.703.092.849.006.303.058.794.026
Text (BERT base ).885.025.663.093.838.020.287.022.765.035
Text (BERT large ).922.006.793.018.875.016.357.039.834.002

Binary subset with predicted or oracle evidence

Binary subset: not comparable with the main three-class document test. Oracle uses gold evidence and is not a deployable model result.

Published source

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Binary subset with predicted or oracle evidence. Published source metrics with original precision; unreported cells are not zero.
SystemNLI accuracy meanAccuracy std.Contradiction F1 meanContradiction F1 std.Entailment F1 meanEntailment F1 std.
Majority vote.814Unreported.239Unreported.645Unreported
Span NLI (BERT base ).883.006.490.007.795.005
Span NLI (BERT large ).899.004.492.065.820.012
Oracle NLI (BERT base ).918.005.657.062.816.006
Oracle NLI (BERT large ).908.011.620.082.806.015

Negation-condition ablation

Paper ablation; numeric cells remain in their source units and cannot be pooled with the main test.

Published source

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Negation-condition ablation. Published source metrics with original precision; unreported cells are not zero.
ConditionMajority-label NLI accuracyMinority-label accuracyWeighted accuracyMinority-label share (%)
w/o (local).91.77.8421
w/ (local).92.40.667
w/o (non-local).98.72.8519
w/ (non-local).90.00.456

Continuous and discontinuous evidence

Paper ablation; numeric cells remain in their source units and cannot be pooled with the main test.

Published source

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Continuous and discontinuous evidence. Published source metrics with original precision; unreported cells are not zero.
Evidence shapeContext nMean spansSpans before one evidence spanSpans before all evidenceEvidence mAP
Continuous1282.641.093.820.91
Discontinuous1282.341.043.840.94
Continuous642.641.164.330.89
Discontinuous642.341.014.850.94

Reference-condition ablation

Paper ablation; numeric cells remain in their source units and cannot be pooled with the main test.

Published source

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. CC BY 4.0. Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Reference-condition ablation. Published source metrics with original precision; unreported cells are not zero.
ConditionMajority-label NLI accuracyMinority-label accuracyWeighted accuracyMinority-label share (%)
w/o Reference.91.88.8926
w/ Reference.93——0

Original benchmark results and configurations

ContractNLI evaluates both document-level three-class inference and evidence retrieval. Accuracy, F1, mAP, and P@R80 are 0–1 ratios. Means and standard deviations occupy separate columns. Oracle evidence experiments use a distinct binary subset. We preserve the paper’s differing BERT-large text-hypothesis accuracy deviations rather than choosing one.

Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.

Snapshot

7 source tables19 configurationsDisplay only

About ContractNLI

Year

2021

Tasks

Contract entailment and evidence identification

Format

Decision classification

Difficulty

Depends on the evaluated split and label mapping

ContractNLI evaluates both document-level three-class inference and evidence retrieval. Accuracy, F1, mAP, and P@R80 are 0–1 ratios. Means and standard deviations occupy separate columns. Oracle evidence experiments use a distinct binary subset. We preserve the paper’s differing BERT-large text-hypothesis accuracy deviations rather than choosing one.

Freshness and provenance

Version

ContractNLI 2021

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

Which ContractNLI results are included?

This page includes 7 numeric result and configuration tables from ContractNLI’s original paper and available benchmark-owner updates, covering 19 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

Can I compare these ContractNLI scores with the Perplexity panel?

Compare ContractNLI scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

Where can I download the ContractNLI results?

The results download on this page provides every imported ContractNLI table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.

Last updated: October 1, 2026 source review · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.