# DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) vs Qwen3.8 Max

> Side-by-side benchmark comparison for DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) and Qwen3.8 Max across agentic, coding, multimodal, knowledge, reasoning, multilingual, and math tasks.

- Canonical page: https://benchlm.ai/compare/native-contractnli-deberta-v2-xlarge-fine-tuned-on-span-identification-hendrycks-et-al-2021-vs-qwen3-8-max
- Last updated: October 1, 2026

- Shared sourced benchmarks: 0
- HTML indexing: indexable
- Ranking lane: BenchAlign v5.8

## Quick Verdict

No benchmark verdict yet: there is no shared sourced benchmark row for DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) and Qwen3.8 Max.

## Summary

- DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) and Qwen3.8 Max do not currently share a sourced benchmark row. The mirror therefore compares metadata, pricing, context windows, and each model's separately reported coverage without naming a benchmark winner.

## Model Snapshot

| Property | DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) | Qwen3.8 Max |
|----------|----------|----------|
| Creator | Koreeda and Manning | Alibaba |
| Type | Research system | Open Weight |
| Reasoning | Unspecified | Reasoning |
| Context | null | 1M |
| Independent public score (not head-to-head) | null | 72.1 |
| Benchmarks Covered | 0 | 60 |
| Pricing (input/output) | Pricing unavailable / Pricing unavailable | Pricing unavailable / Pricing unavailable |

## Category Breakdown

### Agentic

- Winner: Insufficient shared evidence
- DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) public-lane score: Coming soon
- Qwen3.8 Max public-lane score: 65.4 (Supported · #14/119)

| Benchmark | DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) | Qwen3.8 Max | Winner |
|-----------|-----------|-----------|--------|
| Terminal-Bench 2.1 | Coming soon | 86.6% | Coming soon |
| CoWorkBench | Coming soon | 74.8% | Coming soon |
| JobBench | Coming soon | 53.4% | Coming soon |
| skillsBench | Coming soon | 70.2% | Coming soon |
| Agents' Last Exam | Coming soon | 52.4% | Coming soon |
| AutomationBench | Coming soon | 27.3% | Coming soon |
| Toolathlon-Verified | Coming soon | 72.5% | Coming soon |
| WideResearch | Coming soon | 81.9% | Coming soon |
| HLE w/ tools | Coming soon | 56.2% | Coming soon |
| OSWorld-Verified | Coming soon | 86.1% | Coming soon |
| OSWorld 2.0 | Coming soon | 19.4% | Coming soon |
| WebArena-Verified | Coming soon | 66.8% | Coming soon |
| AndroidWorld | Coming soon | 85.3% | Coming soon |
| MobileWorld | Coming soon | 77.8% | Coming soon |
| Terminal-Bench 2.1 (Vals) | Coming soon | 67.4% | Coming soon |

### Coding

- Winner: Insufficient shared evidence
- DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) public-lane score: Coming soon
- Qwen3.8 Max public-lane score: 55.4 (Supported · #30/144)

| Benchmark | DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) | Qwen3.8 Max | Winner |
|-----------|-----------|-----------|--------|
| Terminal-Bench 2.1 | Coming soon | 86.6% | Coming soon |
| SWE-bench Pro | Coming soon | 67.7% | Coming soon |
| DeepSWE | Coming soon | 56.6% | Coming soon |
| NL2Repo | Coming soon | 55.9% | Coming soon |
| FrontierSWE | Coming soon | 73.5% | Coming soon |
| MLS-Bench Lite | Coming soon | 41.0% | Coming soon |
| PaperBench | Coming soon | 93.0% | Coming soon |
| VulcanBench v3 | Coming soon | 81.2% | Coming soon |
| OpenHarmony Bench | Coming soon | 60.8% | Coming soon |
| FrontierSWE v2 | Coming soon | 15.8% | Coming soon |
| LiveCodeBench (Vals) | Coming soon | 87.9% | Coming soon |
| SWE-bench (Vals) | Coming soon | 85.6% | Coming soon |

### Reasoning

- Winner: Insufficient shared evidence
- DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) public-lane score: Coming soon
- Qwen3.8 Max public-lane score: 87.7 (Unranked · 2 rankable rows)

| Benchmark | DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) | Qwen3.8 Max | Winner |
|-----------|-----------|-----------|--------|
| MRCRv2 | Coming soon | 92.9% | Coming soon |
| LongBench v2 | Coming soon | 66.3% | Coming soon |

### Multimodal & Grounded

- Winner: Insufficient shared evidence
- DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) public-lane score: Coming soon
- Qwen3.8 Max public-lane score: 88.4 (#5/49)

| Benchmark | DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) | Qwen3.8 Max | Winner |
|-----------|-----------|-----------|--------|
| MMMU-Pro | Coming soon | 82.3% | Coming soon |
| MathVision | Coming soon | 95.2% | Coming soon |
| MathVision w/ Python | Coming soon | 97.7% | Coming soon |
| BabyVision | Coming soon | 82.0% | Coming soon |
| BabyVision w/ Python | Coming soon | 91.3% | Coming soon |
| ZeroBench | Coming soon | 24.0% | Coming soon |
| ZeroBench w/ Python | Coming soon | 49.0% | Coming soon |
| MedXpertQA (MM) | Coming soon | 80.4% | Coming soon |
| ScreenSpot Pro | Coming soon | 84.5% | Coming soon |
| Vision2Web | Coming soon | 69.0% | Coming soon |
| CharXiv w/o tools | Coming soon | 88.4% | Coming soon |
| CharXiv | Coming soon | 93.5% | Coming soon |
| OmniDocBench 1.5 | Coming soon | 92.1% | Coming soon |
| OCRBench V2 | Coming soon | 74.2% | Coming soon |
| CC-OCR | Coming soon | 79.6% | Coming soon |
| RealWorldQA | Coming soon | 88.0% | Coming soon |
| ERQA | Coming soon | 77.8% | Coming soon |
| SimpleVQA | Coming soon | 75.0% | Coming soon |
| PerceptionBench | Coming soon | 63.5% | Coming soon |
| Video-MME (with subtitle) | Coming soon | 90.4% | Coming soon |
| VideoMMMU | Coming soon | 88.7% | Coming soon |
| MMVU | Coming soon | 82.4% | Coming soon |
| MLVU (M-Avg) | Coming soon | 90.8% | Coming soon |
| LVBench | Coming soon | 81.8% | Coming soon |

### Knowledge

- Winner: Insufficient shared evidence
- DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) public-lane score: Coming soon
- Qwen3.8 Max public-lane score: 66.1 (Supported · #23/170)

| Benchmark | DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) | Qwen3.8 Max | Winner |
|-----------|-----------|-----------|--------|
| GPQA | Coming soon | 92.6% | Coming soon |
| GPQA-D | Coming soon | 92.6% | Coming soon |
| HLE | Coming soon | 43.6% | Coming soon |
| HLE w/o tools | Coming soon | 43.6% | Coming soon |
| GPQA Diamond (Vals) | Coming soon | 93.7% | Coming soon |
| MMLU-Pro (Vals) | Coming soon | 88.6% | Coming soon |

### Instruction Following

- Winner: Insufficient shared evidence
- DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) public-lane score: Coming soon
- Qwen3.8 Max public-lane score: 90.5 (#16/124)

| Benchmark | DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) | Qwen3.8 Max | Winner |
|-----------|-----------|-----------|--------|
| IFBench | Coming soon | 82.8% | Coming soon |

## FAQ

### Can I compare DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) and Qwen3.8 Max on BenchLM yet?

Not fully yet. BenchLM is tracking both models, but sourced benchmark coverage is still incomplete for a fair score-level comparison.

### Why does this page show "coming soon" values?

BenchLM only calls winners when public benchmark coverage is available on both sides of the comparison.

## Related Comparisons

- [DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) vs Qwen3.8 Max Preview](/compare/native-contractnli-deberta-v2-xlarge-fine-tuned-on-span-identification-hendrycks-et-al-2021-vs-qwen3-8-max-preview)
- [Qwen3.8 Max vs Qwen3.8 Max Preview](/compare/qwen3-8-max-vs-qwen3-8-max-preview)
- [DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) vs GPT-6 Astra](/compare/gpt-6-astra-vs-native-contractnli-deberta-v2-xlarge-fine-tuned-on-span-identification-hendrycks-et-al-2021)
- [Qwen3.8 Max vs GPT-6 Astra](/compare/gpt-6-astra-vs-qwen3-8-max)

## Explore More

- [DeBERTa v2 xlarge / Fine-tuned on span identification Hendrycks et al. (2021) (ContractNLI) profile](/models/native-contractnli-deberta-v2-xlarge-fine-tuned-on-span-identification-hendrycks-et-al-2021)
- [Qwen3.8 Max profile](/models/qwen3-8-max)
- [Compare Pricing](/llm-pricing)
- [Alternative Finder](/tools/alternative-finder)
- [LLM Selector](/tools/llm-selector)
- [Overall Rankings](/best/overall)
