# Qwen3.7 Plus Benchmark Scores & Performance

> Qwen3.7 Plus by Alibaba scores 62.29/100 overall, ranking #57 out of 411 AI models.

Canonical page: https://benchlm.ai/models/qwen3-7-plus

Last updated: September 4, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Alibaba |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Overall Score | 62.29/100 |
| Overall Rank | #57 of 411 |

## Family & Coverage

- Family: Qwen3.7 Plus
- Variant: base
- Benchmarks covered: 70 of 422
- Related earlier model: [Qwen3.6 Plus](/models/qwen3-6-plus)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 70.3% |
| [QwenClawBench](/benchmarks/qwenclawbench) | 61.8% |
| [QwenWebBench](/benchmarks/qwenwebbench) | 1536 |
| [Claw-Eval](/benchmarks/claw-eval) | 62.7% |
| [BFCL v4](/benchmarks/bfcl-v4) | 72.9% |
| [MCP Atlas](/benchmarks/mcpatlas) | 73.2% |
| [VITA-Bench](/benchmarks/vitabench) | 45.6% |
| [DeepPlanning](/benchmarks/deepplanning) | 62.3% |
| [OSWorld-Verified](/benchmarks/osworld-verified) | 73.3% |
| [AndroidWorld](/benchmarks/androidworld) | 81.0% |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 20.7% |
| [APEX-Agents-AA](/benchmarks/apexagentsaa) | 22.4% |
| [τ²-bench results](/benchmarks/tau2-bench) | 93% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 22.4% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 947 |
| [OSWorld 2.0](/benchmarks/osworld2) | 2.8% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 52.8% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 70.3% |
| [SWE-bench Verified](/benchmarks/swe-bench-verified) | 77.7% |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 57.6% |
| [SWE Multilingual](/benchmarks/swe-bench-multilingual) | 75.8% |
| [NL2Repo](/benchmarks/nl2repo) | 41.1% |
| [SciCode](/benchmarks/scicode) | 51.3% |
| [LiveCodeBench](/benchmarks/livecodebench) | 89.6% |
| [AA Coding Index](/benchmarks/aacodingindex) | 55.9% |
| [AA-SciCode](/benchmarks/aascicode) | 45.5% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMMU-Pro](/benchmarks/mmmu-pro) | 79% |
| [MathVision](/benchmarks/mathvision) | 90.3% |
| [CharXiv](/benchmarks/charxiv) | 85.9% |
| [ERQA](/benchmarks/erqa) | 69.8% |
| [MedXpertQA (MM)](/benchmarks/medxpertqamm) | 71.0% |
| [ScreenSpot Pro](/benchmarks/screenspot-pro) | 79.0% |
| [SimpleVQA](/benchmarks/simplevqa) | 81.7% |
| [MMSearch-Plus](/benchmarks/mmsearchplus) | 41.4% |
| [RealWorldQA](/benchmarks/realworldqa) | 86.9% |
| [OmniDocBench 1.5](/benchmarks/omnidocbench15) | 91.4% |
| [OCRBench V2](/benchmarks/ocrbenchv2) | 70.7% |
| [ODINW13](/benchmarks/odinw13) | 51.1% |
| [Video-MME (with subtitle)](/benchmarks/videommewithsub) | 88.0% |
| [VideoMMMU](/benchmarks/videommmu) | 85.4% |
| [MLVU (M-Avg)](/benchmarks/mlvuavg) | 87.4% |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 80.5% |
| [Design Arena Website](/benchmarks/designarenawebsite) | 1284 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [CritPt](/benchmarks/critpt) | 9.1% |
| [MRCRv2](/benchmarks/mrcrv2) | 91.7% |
| [AA-LCR](/benchmarks/lcr) | 69.0% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA](/benchmarks/gpqa) | 90.3% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 90.3% |
| [HLE](/benchmarks/hle) | 34.7% |
| [MMLU-Pro](/benchmarks/mmlu-pro) | 88.5% |
| [MMLU-Redux](/benchmarks/mmluredux) | 94.5% |
| [SuperGPQA](/benchmarks/supergpqa) | 71.4% |
| [MMMLU](/benchmarks/mmmlu) | 89.0% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 39.4% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 90.0% |
| [AA-HLE](/benchmarks/aahle) | 35.6% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 1.1% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 22.5% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 27.7% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [IFEval](/benchmarks/ifeval) | 94.6% |
| [IFBench](/benchmarks/ifbench) | 79.1% |
| [AA-IFBench](/benchmarks/aaifbench) | 78.0% |

## Multilingual Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMLU-ProX](/benchmarks/mmluprox) | 85.4% |
| [NOVA-63](/benchmarks/nova63) | 58.8% |
| [INCLUDE](/benchmarks/include) | 83.0% |
| [MAXIFE](/benchmarks/maxife) | 88.8% |
| [PolyMath](/benchmarks/polymath) | 84.0% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [HMMT Feb 2026](/benchmarks/hmmtfeb2026) | 92.9% |
| [IMOAnswerBench](/benchmarks/imoanswerbench) | 86.0% |
| [Apex](/benchmarks/apex) | 22.7% |

## Other Alibaba Models

- [Qwen3.8 Max](/models/qwen3-8-max) - Score: 72.43
- [Qwen3.7 Max](/models/qwen3-7-max) - Score: 68.56
- [Qwen3.8-27B](/models/qwen3-8-27b) - Score: 68.35
- [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) - Score: 63.64
- [Qwen3.6 Plus](/models/qwen3-6-plus) - Score: 61.49
- [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) - Score: 59.42
- [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) - Score: 59.01
- [Qwen3.5 397B (Reasoning)](/models/qwen3-5-397b-reasoning) - Score: 58.89
- [Qwen3.5-27B](/models/qwen3-5-27b) - Score: 58.31
- [Qwen3 235B 2507 (Reasoning)](/models/qwen3-235b-2507-reasoning) - Score: 57.43
- [Qwen3.5 397B](/models/qwen3-5-397b) - Score: 56.44
- [Qwen3.5 Flash](/models/qwen3-5-flash) - Score: 56.03
- [Qwen3 235B 2507](/models/qwen3-235b-2507) - Score: 55.46
- [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) - Score: 54.86
- [Qwen3.6-27B](/models/qwen3-6-27b) - Score: 52.72
- [Qwen3.5 Plus](/models/qwen3-5-plus) - Score: 51.63
- [Qwen3.7 Flash](/models/qwen3-7-flash) - Score: 50.75
- [Qwen2.5-1M](/models/qwen2-5-1m) - Score: 49.45
- [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) - Score: 48.34
- [Qwen3 Max](/models/qwen3-max) - Score: 43.82
- [Qwen2.5-VL-32B](/models/qwen2-5-vl-32b) - Score: 39.65
- [Qwen2.5-72B](/models/qwen2-5-72b) - Score: 37.32
- [Qwen2.5 Coder 32B Instruct](/models/qwen2-5-coder-32b-instruct) - Score: 34.02
- [Qwen3.8 Max Preview](/models/qwen3-8-max-preview) - Score: not computed
- [Qwen2.5-Omni 7B](/models/qwen2-5-omni-7b) - Score: not computed
- [Qwen2-Audio 7B Instruct](/models/qwen2-audio-7b-instruct) - Score: not computed
- [Qwen-Audio 7B](/models/qwen-audio-7b) - Score: not computed
- [Qwen-Audio-Chat 7B](/models/qwen-audio-chat-7b) - Score: not computed
