# Qwen3.8-27B Benchmark Scores & Performance

> Qwen3.8-27B by Alibaba scores 55.63/100 overall, ranking #55 out of 505 AI models.

Canonical page: https://benchlm.ai/models/qwen3-8-27b

Last updated: September 22, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Alibaba |
| Source Type | Open Weight |
| Reasoning Type | Reasoning |
| Context Window | 262K |
| Official model card | [Qwen3.8-27B model card](https://huggingface.co/Qwen/Qwen3.8-27B) |
| Overall Score | 55.63/100 |
| Overall Rank | #55 of 505 |

## Family & Coverage

- Family: Qwen3.8-27B
- Variant: base
- Benchmarks covered: 50 of 481
- Related earlier model: [Qwen3.6-27B](/models/qwen3-6-27b)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 73.0% |
| [CoWorkBench](/benchmarks/coworkbench) | 70.7% |
| [JobBench](/benchmarks/jobbench) | 33.4% |
| [Agents' Last Exam](/benchmarks/agentslastexam) | 42.9% |
| [OSWorld-Verified](/benchmarks/osworld-verified) | 84.3% |
| [WebArena-Verified](/benchmarks/webarena-verified) | 64.8% |
| [AndroidWorld](/benchmarks/androidworld) | 81.9% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 45.4% |
| [AA Tau3 Banking](/benchmarks/aatau3banking) | 48.0% |
| [AA EnterpriseOps-Gym](/benchmarks/aaenterpriseopsgym) | 44.2% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 58.4% |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 46.5% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1463 |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 73.0% |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 61.7% |
| [NL2Repo](/benchmarks/nl2repo) | 42.3% |
| [DeepSWE](/benchmarks/deepswe) | 42.2% |
| [LiveCodeBench v6](/benchmarks/livecodebench-v6) | 90.3% |
| [AA-SciCode](/benchmarks/aascicode) | 46.6% |
| [VulcanBench v3](/benchmarks/vulcanbench) | 82.6% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 84.0% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 86.0% |
| [AA Coding Index](/benchmarks/aacodingindex) | 68.1% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MathVision](/benchmarks/mathvision) | 90.0% |
| [MathVision w/ Python](/benchmarks/mathvisionpython) | 94.6% |
| [BabyVision](/benchmarks/babyvision) | 65.7% |
| [BabyVision w/ Python](/benchmarks/babyvisionpython) | 85.6% |
| [Vision2Web](/benchmarks/vision2web) | 62.9% |
| [CharXiv w/o tools](/benchmarks/charxivnotools) | 83.7% |
| [CharXiv](/benchmarks/charxiv) | 90.2% |
| [OmniDocBench 1.5](/benchmarks/omnidocbench15) | 91.1% |
| [RealWorldQA](/benchmarks/realworldqa) | 85.9% |
| [ERQA](/benchmarks/erqa) | 65.5% |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 76.3% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 82.0% |
| [CritPt](/benchmarks/critpt) | 5.4% |
| [MLCR-AA](/benchmarks/aamlcr) | 21.7% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA](/benchmarks/gpqa) | 89.2% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 89.2% |
| [HLE](/benchmarks/hle) | 30.8% |
| [HLE w/o tools](/benchmarks/hlenotools) | 30.8% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 33.7% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 90.5% |
| [AA-HLE](/benchmarks/aahle) | 33.9% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | -10.0% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 15.6% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 30.3% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 88.9% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 84.3% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [IFBench](/benchmarks/ifbench) | 79.5% |

## Other Alibaba Models

- [Qwen3.8 Max](/models/qwen3-8-max) - Score: 72.02
- [Qwen3.7 Max](/models/qwen3-7-max) - Score: 62.89
- [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) - Score: 60.67
- [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) - Score: 56.07
- [Qwen3.7 Plus](/models/qwen3-7-plus) - Score: 55.76
- [Qwen3.6 Plus](/models/qwen3-6-plus) - Score: 55.22
- [Qwen3.5 Plus](/models/qwen3-5-plus) - Score: 48.78
- [Qwen3.7 Flash](/models/qwen3-7-flash) - Score: 47.44
- [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) - Score: 47.02
- [Qwen3.6-27B](/models/qwen3-6-27b) - Score: 46.27
- [Qwen3.5 Flash](/models/qwen3-5-flash) - Score: 45.42
- [Qwen3.5-27B](/models/qwen3-5-27b) - Score: 43.82
- [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) - Score: 41.55
- [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) - Score: 40.06
- [Qwen3 Max](/models/qwen3-max) - Score: 39.09
- [Qwen2.5-72B](/models/qwen2-5-72b) - Score: 31.32
- [Qwen3-Omni-30B-A3B-Instruct](/models/qwen3-omni-30b-a3b-instruct) - Score: 31.05
- [Qwen2.5 Coder 32B Instruct](/models/qwen2-5-coder-32b-instruct) - Score: 27.1
- [Qwen3.8-Omni-Flash](/models/qwen3-8-omni-flash) - Score: not computed
- [Qwen3 235B 2507](/models/qwen3-235b-2507) - Score: not computed
- [Qwen3.5 397B](/models/qwen3-5-397b) - Score: not computed
- [Qwen3 235B 2507 (Reasoning)](/models/qwen3-235b-2507-reasoning) - Score: not computed
- [Qwen3.5 397B (Reasoning)](/models/qwen3-5-397b-reasoning) - Score: not computed
- [Qwen2.5-1M](/models/qwen2-5-1m) - Score: not computed
- [Qwen2.5-VL-32B](/models/qwen2-5-vl-32b) - Score: not computed
- [Qwen3.8 Max Preview](/models/qwen3-8-max-preview) - Score: not computed
- [Qwen3-ASR 0.6B](/models/qwen3-asr-0-6b) - Score: not computed
- [Qwen3-ASR 1.7B](/models/qwen3-asr-1-7b) - Score: not computed
- [Qwen-AgentWorld-35B-A3B](/models/qwen-agentworld-35b-a3b) - Score: not computed
- [Qwen3-Omni-30B-A3B-Thinking](/models/qwen3-omni-30b-a3b-thinking) - Score: not computed
- [Qwen2.5-Omni 7B](/models/qwen2-5-omni-7b) - Score: not computed
- [Qwen2-Audio 7B Instruct](/models/qwen2-audio-7b-instruct) - Score: not computed
- [Qwen-Audio 7B](/models/qwen-audio-7b) - Score: not computed
- [Qwen-Audio-Chat 7B](/models/qwen-audio-chat-7b) - Score: not computed
- [Qwen4 27B](/models/qwen4-27b) - Score: not computed
- [Qwen4 Flash & Plus](/models/qwen4-flash-plus) - Score: not computed
- [Qwen4 Max](/models/qwen4-max) - Score: not computed
