# Qwen3.6-27B Benchmark Scores & Performance

> Qwen3.6-27B by Alibaba scores 52.73/100 overall, ranking #113 out of 411 AI models.

Canonical page: https://benchlm.ai/models/qwen3-6-27b

Last updated: September 4, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Alibaba |
| Source Type | Open Weight |
| Reasoning Type | Reasoning |
| Context Window | 262K |
| Overall Score | 52.73/100 |
| Overall Rank | #113 of 411 |

## Family & Coverage

- Family: Qwen3.6-27B
- Variant: base
- Benchmarks covered: 54 of 422
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 59.3% |
| [Claw-Eval](/benchmarks/claw-eval) | 72.4% |
| [QwenClawBench](/benchmarks/qwenclawbench) | 53.4% |
| [QwenWebBench](/benchmarks/qwenwebbench) | 1487 |
| [AndroidWorld](/benchmarks/androidworld) | 70.3% |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 27.5% |
| [τ²-bench results](/benchmarks/tau2-bench) | 94.2% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 31.9% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1138 |
| [Gert Labs](/benchmarks/gertlabs) | 54.84% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [SWE-bench Verified](/benchmarks/swe-bench-verified) | 77.2% |
| [SWE Multilingual](/benchmarks/swe-bench-multilingual) | 71.3% |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 53.5% |
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 59.3% |
| [LiveCodeBench](/benchmarks/livecodebench) | 83.9% |
| [NL2Repo](/benchmarks/nl2repo) | 36.2% |
| [AA Coding Index](/benchmarks/aacodingindex) | 53.7% |
| [AA-SciCode](/benchmarks/aascicode) | 39.8% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMMU](/benchmarks/mmmu) | 82.9% |
| [MMMU-Pro](/benchmarks/mmmu-pro) | 75.8% |
| [RealWorldQA](/benchmarks/realworldqa) | 84.1% |
| [DynaMath](/benchmarks/dynamath) | 85.6% |
| [MStar](/benchmarks/mstar) | 81.4% |
| [SimpleVQA](/benchmarks/simplevqa) | 56.1% |
| [CharXiv](/benchmarks/charxiv) | 78.4% |
| [CC-OCR](/benchmarks/ccocr) | 81.2% |
| [CountBench](/benchmarks/countbench) | 97.8% |
| [RefCOCO (avg)](/benchmarks/refcocoavg) | 92.5% |
| [ERQA](/benchmarks/erqa) | 62.5% |
| [Video-MME (with subtitle)](/benchmarks/videommewithsub) | 87.7% |
| [VideoMMMU](/benchmarks/videommmu) | 84.4% |
| [MLVU (M-Avg)](/benchmarks/mlvuavg) | 86.6% |
| [V*](/benchmarks/vstar) | 94.7% |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 74.6% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 73.3% |
| [CritPt](/benchmarks/critpt) | 1.1% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMLU-Pro](/benchmarks/mmlu-pro) | 86.2% |
| [MMLU-Redux](/benchmarks/mmluredux) | 93.5% |
| [SuperGPQA](/benchmarks/supergpqa) | 66% |
| [C-Eval](/benchmarks/ceval) | 91.4% |
| [GPQA](/benchmarks/gpqa) | 87.8% |
| [HLE](/benchmarks/hle) | 24% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 37.7% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 84.2% |
| [AA-HLE](/benchmarks/aahle) | 23.1% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | -20.0% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 19.6% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 49.3% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-IFBench](/benchmarks/aaifbench) | 67.6% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [HMMT Feb 2025](/benchmarks/hmmtfeb2025) | 93.8% |
| [HMMT Nov 2025](/benchmarks/hmmtnov2025) | 90.7% |
| [HMMT Feb 2026](/benchmarks/hmmtfeb2026) | 84.3% |
| [MMAnswerBench](/benchmarks/mmanswerbench) | 80.8% |
| [AIME26](/benchmarks/aime2026) | 94.1% |

## Other Alibaba Models

- [Qwen3.8 Max](/models/qwen3-8-max) - Score: 72.43
- [Qwen3.7 Max](/models/qwen3-7-max) - Score: 68.56
- [Qwen3.8-27B](/models/qwen3-8-27b) - Score: 68.35
- [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) - Score: 63.64
- [Qwen3.7 Plus](/models/qwen3-7-plus) - Score: 62.3
- [Qwen3.6 Plus](/models/qwen3-6-plus) - Score: 61.49
- [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) - Score: 59.43
- [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) - Score: 59.01
- [Qwen3.5 397B (Reasoning)](/models/qwen3-5-397b-reasoning) - Score: 58.89
- [Qwen3.5-27B](/models/qwen3-5-27b) - Score: 58.32
- [Qwen3 235B 2507 (Reasoning)](/models/qwen3-235b-2507-reasoning) - Score: 57.43
- [Qwen3.5 397B](/models/qwen3-5-397b) - Score: 56.44
- [Qwen3.5 Flash](/models/qwen3-5-flash) - Score: 56.03
- [Qwen3 235B 2507](/models/qwen3-235b-2507) - Score: 55.47
- [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) - Score: 54.86
- [Qwen3.5 Plus](/models/qwen3-5-plus) - Score: 51.64
- [Qwen3.7 Flash](/models/qwen3-7-flash) - Score: 50.76
- [Qwen2.5-1M](/models/qwen2-5-1m) - Score: 49.45
- [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) - Score: 48.34
- [Qwen3 Max](/models/qwen3-max) - Score: 43.82
- [Qwen2.5-VL-32B](/models/qwen2-5-vl-32b) - Score: 39.65
- [Qwen2.5-72B](/models/qwen2-5-72b) - Score: 37.32
- [Qwen2.5 Coder 32B Instruct](/models/qwen2-5-coder-32b-instruct) - Score: 34.03
- [Qwen3.8 Max Preview](/models/qwen3-8-max-preview) - Score: not computed
- [Qwen2.5-Omni 7B](/models/qwen2-5-omni-7b) - Score: not computed
- [Qwen2-Audio 7B Instruct](/models/qwen2-audio-7b-instruct) - Score: not computed
- [Qwen-Audio 7B](/models/qwen-audio-7b) - Score: not computed
- [Qwen-Audio-Chat 7B](/models/qwen-audio-chat-7b) - Score: not computed
