# Qwen3.8 Max Benchmark Scores & Performance

> Qwen3.8 Max by Alibaba scores 71.76/100 overall, ranking #10 out of 486 AI models.

Canonical page: https://benchlm.ai/models/qwen3-8-max

Last updated: September 15, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Alibaba |
| Source Type | Open Weight |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Official model card | [Qwen3.8-2.4T-A95B model card](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B) |
| Overall Score | 71.76/100 |
| Overall Rank | #10 of 486 |

## Family & Coverage

- Family: Qwen3.8 Max
- Variant: base
- Benchmarks covered: 61 of 439
- Sibling models: [Qwen3.8 Max Preview](/models/qwen3-8-max-preview)
- Related earlier model: [Qwen3.8 Max Preview](/models/qwen3-8-max-preview)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 86.6% |
| [CoWorkBench](/benchmarks/coworkbench) | 74.8% |
| [JobBench](/benchmarks/jobbench) | 53.4% |
| [skillsBench](/benchmarks/skillsbench) | 70.2% |
| [Agents' Last Exam](/benchmarks/agentslastexam) | 52.4% |
| [AutomationBench](/benchmarks/automationbench) | 27.3% |
| [Toolathlon-Verified](/benchmarks/toolathlonverified) | 72.5% |
| [WideResearch](/benchmarks/wideresearch) | 81.9% |
| [HLE w/ tools](/benchmarks/hlewithtools) | 56.2% |
| [OSWorld-Verified](/benchmarks/osworld-verified) | 86.1% |
| [OSWorld 2.0](/benchmarks/osworld2) | 19.4% |
| [WebArena-Verified](/benchmarks/webarena-verified) | 66.8% |
| [AndroidWorld](/benchmarks/androidworld) | 85.3% |
| [MobileWorld](/benchmarks/mobileworld) | 77.8% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 67.4% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 86.6% |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 67.7% |
| [DeepSWE](/benchmarks/deepswe) | 56.6% |
| [NL2Repo](/benchmarks/nl2repo) | 55.9% |
| [FrontierSWE](/benchmarks/frontierswe) | 73.5% |
| [MLS-Bench Lite](/benchmarks/mlsbenchlite) | 41.0% |
| [PaperBench](/benchmarks/paperbench) | 93.0% |
| [QwenReactBench](/benchmarks/qwenreactbench) | 1724 |
| [VulcanBench v3](/benchmarks/vulcanbench) | 81.2% |
| [OpenHarmony Bench](/benchmarks/openharmonybench) | 60.8% |
| [FrontierSWE v2](/benchmarks/frontierswev2) | 15.8% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 87.9% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 85.6% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMMU-Pro](/benchmarks/mmmu-pro) | 82.3% |
| [MathVision](/benchmarks/mathvision) | 95.2% |
| [MathVision w/ Python](/benchmarks/mathvisionpython) | 97.7% |
| [BabyVision](/benchmarks/babyvision) | 82.0% |
| [BabyVision w/ Python](/benchmarks/babyvisionpython) | 91.3% |
| [ZeroBench](/benchmarks/zerobench) | 24.0% |
| [ZeroBench w/ Python](/benchmarks/zerobenchpython) | 49.0% |
| [MedXpertQA (MM)](/benchmarks/medxpertqamm) | 80.4% |
| [ScreenSpot Pro](/benchmarks/screenspot-pro) | 84.5% |
| [Vision2Web](/benchmarks/vision2web) | 69.0% |
| [CharXiv w/o tools](/benchmarks/charxivnotools) | 88.4% |
| [CharXiv](/benchmarks/charxiv) | 93.5% |
| [OmniDocBench 1.5](/benchmarks/omnidocbench15) | 92.1% |
| [OCRBench V2](/benchmarks/ocrbenchv2) | 74.2% |
| [CC-OCR](/benchmarks/ccocr) | 79.6% |
| [RealWorldQA](/benchmarks/realworldqa) | 88.0% |
| [ERQA](/benchmarks/erqa) | 77.8% |
| [SimpleVQA](/benchmarks/simplevqa) | 75.0% |
| [PerceptionBench](/benchmarks/perceptionbench) | 63.5% |
| [Video-MME (with subtitle)](/benchmarks/videommewithsub) | 90.4% |
| [VideoMMMU](/benchmarks/videommmu) | 88.7% |
| [MMVU](/benchmarks/mmvu) | 82.4% |
| [MLVU (M-Avg)](/benchmarks/mlvuavg) | 90.8% |
| [LVBench](/benchmarks/lvbench) | 81.8% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MRCRv2](/benchmarks/mrcrv2) | 92.9% |
| [LongBench v2](/benchmarks/longbench-v2) | 66.3% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA](/benchmarks/gpqa) | 92.6% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 92.6% |
| [HLE](/benchmarks/hle) | 43.6% |
| [HLE w/o tools](/benchmarks/hlenotools) | 43.6% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 93.7% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 88.6% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [IFBench](/benchmarks/ifbench) | 82.8% |

## Other Alibaba Models

- [Qwen3.7 Max](/models/qwen3-7-max) - Score: 67.03
- [Qwen3.8-27B](/models/qwen3-8-27b) - Score: 64.65
- [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) - Score: 63.83
- [Qwen3.7 Plus](/models/qwen3-7-plus) - Score: 61.78
- [Qwen3.6 Plus](/models/qwen3-6-plus) - Score: 60.48
- [Qwen3.5 397B (Reasoning)](/models/qwen3-5-397b-reasoning) - Score: 57.31
- [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) - Score: 57.02
- [Qwen3.5 Flash](/models/qwen3-5-flash) - Score: 56.2
- [Qwen3 235B 2507 (Reasoning)](/models/qwen3-235b-2507-reasoning) - Score: 55.85
- [Qwen3.5-27B](/models/qwen3-5-27b) - Score: 55.38
- [Qwen3.5 397B](/models/qwen3-5-397b) - Score: 54.87
- [Qwen3 235B 2507](/models/qwen3-235b-2507) - Score: 53.89
- [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) - Score: 53.57
- [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) - Score: 50.86
- [Qwen3.5 Plus](/models/qwen3-5-plus) - Score: 50.11
- [Qwen3.7 Flash](/models/qwen3-7-flash) - Score: 49.24
- [Qwen3.6-27B](/models/qwen3-6-27b) - Score: 47.89
- [Qwen2.5-1M](/models/qwen2-5-1m) - Score: 47.88
- [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) - Score: 43.99
- [Qwen3 Max](/models/qwen3-max) - Score: 40.87
- [Qwen3-Omni-30B-A3B-Instruct](/models/qwen3-omni-30b-a3b-instruct) - Score: 38.5
- [Qwen2.5-VL-32B](/models/qwen2-5-vl-32b) - Score: 38.06
- [Qwen2.5-72B](/models/qwen2-5-72b) - Score: 37.55
- [Qwen2.5 Coder 32B Instruct](/models/qwen2-5-coder-32b-instruct) - Score: 33.43
- [Qwen3.8 Max Preview](/models/qwen3-8-max-preview) - Score: not computed
- [Qwen3-ASR 0.6B](/models/qwen3-asr-0-6b) - Score: not computed
- [Qwen3-ASR 1.7B](/models/qwen3-asr-1-7b) - Score: not computed
- [Qwen-AgentWorld-35B-A3B](/models/qwen-agentworld-35b-a3b) - Score: not computed
- [Qwen3-Omni-30B-A3B-Thinking](/models/qwen3-omni-30b-a3b-thinking) - Score: not computed
- [Qwen2.5-Omni 7B](/models/qwen2-5-omni-7b) - Score: not computed
- [Qwen2-Audio 7B Instruct](/models/qwen2-audio-7b-instruct) - Score: not computed
- [Qwen-Audio 7B](/models/qwen-audio-7b) - Score: not computed
- [Qwen-Audio-Chat 7B](/models/qwen-audio-chat-7b) - Score: not computed
