# Span NLI (BERT large ) binary subset (ContractNLI) Benchmark Scores & Performance

> Span NLI (BERT large ) binary subset (ContractNLI) has published results in 1 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-contractnli-span-nli-bert-large-binary-subset

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | Koreeda and Manning |
| Source Type | Research system |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://aclanthology.org/2021.findings-emnlp.164/) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Span NLI (BERT large ) binary subset (ContractNLI)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-contractnli-span-nli-bert-large-binary-subset) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-contractnli-span-nli-bert-large-binary-subset&format=csv)

### Binary subset with predicted or oracle evidence

ContractNLI evaluates both document-level three-class inference and evidence retrieval. Accuracy, F1, mAP, and P@R80 are 0–1 ratios. Means and standard deviations occupy separate columns. Oracle evidence experiments use a distinct binary subset. We preserve the paper’s differing BERT-large text-hypothesis accuracy deviations rather than choosing one.

Binary subset: not comparable with the main three-class document test. Oracle uses gold evidence and is not a deployable model result.

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts — Yuta Koreeda and Christopher D. Manning. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2021 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2021.findings-emnlp.164/) · [Full ContractNLI results](/benchmarks/contractnli)

| System | NLI accuracy mean | Accuracy std. | Contradiction F1 mean | Contradiction F1 std. | Entailment F1 mean | Entailment F1 std. |
| --- | --- | --- | --- | --- | --- | --- |
| Span NLI (BERT large ) | .899 | .004 | .492 | .065 | .820 | .012 |

## Other Koreeda and Manning Models

- [BERT base / Fine-tuned on case law and contract corpora Chalkidis et al. (2020) (ContractNLI)](/models/native-contractnli-bert-base-fine-tuned-on-case-law-and-contract-corpora-chalkidis-et-al-2020) - Score: not computed
- [BERT base / Fine-tuned on NDAs (ContractNLI)](/models/native-contractnli-bert-base-fine-tuned-on-ndas) - Score: not computed
- [BERT base / None (ContractNLI)](/models/native-contractnli-bert-base-none) - Score: not computed
- [BERT base / Pretrained from scratch using a case law corpus Zheng et al. (2021) (ContractNLI)](/models/native-contractnli-bert-base-pretrained-from-scratch-using-a-case-law-corpus-zheng-et-al-2021) - Score: not computed
- [BERT large / Fine-tuned on NDAs (ContractNLI)](/models/native-contractnli-bert-large-fine-tuned-on-ndas) - Score: not computed
- [BERT large / None (ContractNLI)](/models/native-contractnli-bert-large-none) - Score: not computed
- [DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI)](/models/native-contractnli-deberta-v2-xlarge-fine-tuned-on-span-identification-hendrycks-et-al-2021) - Score: not computed
- [DeBERTa v2 xlarge / None (ContractNLI)](/models/native-contractnli-deberta-v2-xlarge-none) - Score: not computed
- [Doc TF-IDF+SVM (ContractNLI)](/models/native-contractnli-doc-tf-idf-svm) - Score: not computed
- [Oracle NLI (BERT base ) binary subset (ContractNLI)](/models/native-contractnli-oracle-nli-bert-base-binary-subset) - Score: not computed
- [Oracle NLI (BERT large ) binary subset (ContractNLI)](/models/native-contractnli-oracle-nli-bert-large-binary-subset) - Score: not computed
- [Span NLI (BERT base ) binary subset (ContractNLI)](/models/native-contractnli-span-nli-bert-base-binary-subset) - Score: not computed
- [Span TF-IDF+Cosine (ContractNLI)](/models/native-contractnli-span-tf-idf-cosine) - Score: not computed
- [Span TF-IDF+SVM (ContractNLI)](/models/native-contractnli-span-tf-idf-svm) - Score: not computed
- [SQuAD (BERT base ) (ContractNLI)](/models/native-contractnli-squad-bert-base) - Score: not computed
- [SQuAD (BERT large ) (ContractNLI)](/models/native-contractnli-squad-bert-large) - Score: not computed
- [Symbol (BERT base ) (ContractNLI)](/models/native-contractnli-symbol-bert-base) - Score: not computed
- [Symbol (BERT large ) (ContractNLI)](/models/native-contractnli-symbol-bert-large) - Score: not computed
