# Qwen3 Max Benchmark Scores & Performance

> Qwen3 Max by Alibaba scores 43.82/100 overall, ranking #166 out of 411 AI models.

Canonical page: https://benchlm.ai/models/qwen3-max

Last updated: September 4, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Alibaba |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Overall Score | 43.82/100 |
| Overall Rank | #166 of 411 |

## Family & Coverage

- Family: Qwen3 Max
- Variant: base
- Benchmarks covered: 14 of 422
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [τ²-bench results](/benchmarks/tau2-bench) | 74.3% |
| [Gert Labs](/benchmarks/gertlabs) | 43.74% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Vibe Code Bench](/benchmarks/vibecodebench) | 3.51% |
| [AA-SciCode](/benchmarks/aascicode) | 38.3% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Design Arena Website](/benchmarks/designarenawebsite) | 1134 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 48.3% |
| [CritPt](/benchmarks/critpt) | 0.0% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 24.4% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 76.4% |
| [AA-HLE](/benchmarks/aahle) | 11.9% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | -43.5% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 24.4% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 89.9% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-IFBench](/benchmarks/aaifbench) | 44.1% |

## Other Alibaba Models

- [Qwen3.8 Max](/models/qwen3-8-max) - Score: 72.43
- [Qwen3.7 Max](/models/qwen3-7-max) - Score: 68.56
- [Qwen3.8-27B](/models/qwen3-8-27b) - Score: 68.35
- [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) - Score: 63.64
- [Qwen3.7 Plus](/models/qwen3-7-plus) - Score: 62.29
- [Qwen3.6 Plus](/models/qwen3-6-plus) - Score: 61.49
- [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) - Score: 59.42
- [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) - Score: 59.01
- [Qwen3.5 397B (Reasoning)](/models/qwen3-5-397b-reasoning) - Score: 58.89
- [Qwen3.5-27B](/models/qwen3-5-27b) - Score: 58.31
- [Qwen3 235B 2507 (Reasoning)](/models/qwen3-235b-2507-reasoning) - Score: 57.43
- [Qwen3.5 397B](/models/qwen3-5-397b) - Score: 56.44
- [Qwen3.5 Flash](/models/qwen3-5-flash) - Score: 56.03
- [Qwen3 235B 2507](/models/qwen3-235b-2507) - Score: 55.46
- [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) - Score: 54.86
- [Qwen3.6-27B](/models/qwen3-6-27b) - Score: 52.72
- [Qwen3.5 Plus](/models/qwen3-5-plus) - Score: 51.63
- [Qwen3.7 Flash](/models/qwen3-7-flash) - Score: 50.75
- [Qwen2.5-1M](/models/qwen2-5-1m) - Score: 49.45
- [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) - Score: 48.34
- [Qwen2.5-VL-32B](/models/qwen2-5-vl-32b) - Score: 39.65
- [Qwen2.5-72B](/models/qwen2-5-72b) - Score: 37.32
- [Qwen2.5 Coder 32B Instruct](/models/qwen2-5-coder-32b-instruct) - Score: 34.02
- [Qwen3.8 Max Preview](/models/qwen3-8-max-preview) - Score: not computed
- [Qwen2.5-Omni 7B](/models/qwen2-5-omni-7b) - Score: not computed
- [Qwen2-Audio 7B Instruct](/models/qwen2-audio-7b-instruct) - Score: not computed
- [Qwen-Audio 7B](/models/qwen-audio-7b) - Score: not computed
- [Qwen-Audio-Chat 7B](/models/qwen-audio-chat-7b) - Score: not computed
