# Qwen3.8-Flash-Next Benchmark Scores & Performance

> Qwen3.8-Flash-Next by Alibaba scores 56.84/100 overall, ranking #77 out of 483 AI models.

Canonical page: https://benchlm.ai/models/qwen3-8-flash-next

Last updated: September 10, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Alibaba |
| Source Type | Open Weight |
| Reasoning Type | Reasoning |
| Context Window | 262K |
| Official model card | [Qwen3.8-Flash-Next model card](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) |
| Overall Score | 56.84/100 |
| Overall Rank | #77 of 483 |

## Family & Coverage

- Family: Qwen3.8-Flash-Next
- Variant: experimental-preview (Next)
- Benchmarks covered: 38 of 435
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [CoWorkBench](/benchmarks/coworkbench) | 73.9% |
| [JobBench](/benchmarks/jobbench) | 55.7% |
| [Agents' Last Exam](/benchmarks/agentslastexam) | 51.2% |
| [Toolathlon-Verified](/benchmarks/toolathlonverified) | 73.5% |
| [AndroidWorld](/benchmarks/androidworld) | 84.5% |
| [OSWorld 2.0](/benchmarks/osworld2) | 19.4% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 57.4% |
| [AA Briefcase](/benchmarks/aabriefcaseelo) | 1587 |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1648 |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 62.5% |
| [SWE Multilingual](/benchmarks/swe-bench-multilingual) | 81% |
| [NL2Repo](/benchmarks/nl2repo) | 48.1% |
| [DeepSWE](/benchmarks/deepswe) | 58.7% |
| [LiveCodeBench v6](/benchmarks/livecodebench-v6) | 91.9% |
| [AA-SciCode](/benchmarks/aascicode) | 50.6% |
| [AA Coding Index](/benchmarks/aacodingindex) | 73.0% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Vision2Web](/benchmarks/vision2web) | 64.0% |
| [ERQA](/benchmarks/erqa) | 72.3% |
| [LVBench](/benchmarks/lvbench) | 76.6% |
| [RealWorldQA](/benchmarks/realworldqa) | 88.5% |
| [MathVision](/benchmarks/mathvision) | 90.6% |
| [MathVision w/ Python](/benchmarks/mathvisionpython) | 95.7% |
| [CharXiv w/o tools](/benchmarks/charxivnotools) | 84.6% |
| [CharXiv](/benchmarks/charxiv) | 90.6% |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 79.8% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 79.7% |
| [CritPt](/benchmarks/critpt) | 11.1% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA](/benchmarks/gpqa) | 91.7% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 91.7% |
| [HLE](/benchmarks/hle) | 35.9% |
| [HLE w/o tools](/benchmarks/hlenotools) | 35.9% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 39.9% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 92.3% |
| [AA-HLE](/benchmarks/aahle) | 38.0% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | -9.7% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 24.5% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 45.3% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [IFBench](/benchmarks/ifbench) | 81.3% |

## Other Alibaba Models

- [Qwen3.8 Max](/models/qwen3-8-max) - Score: 71.63
- [Qwen3.7 Max](/models/qwen3-7-max) - Score: 67.08
- [Qwen3.8-27B](/models/qwen3-8-27b) - Score: 64.52
- [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) - Score: 63.7
- [Qwen3.7 Plus](/models/qwen3-7-plus) - Score: 61.56
- [Qwen3.6 Plus](/models/qwen3-6-plus) - Score: 60.4
- [Qwen3.5 397B (Reasoning)](/models/qwen3-5-397b-reasoning) - Score: 57.15
- [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) - Score: 56.37
- [Qwen3.5 Flash](/models/qwen3-5-flash) - Score: 56.09
- [Qwen3 235B 2507 (Reasoning)](/models/qwen3-235b-2507-reasoning) - Score: 55.69
- [Qwen3.5-27B](/models/qwen3-5-27b) - Score: 55.34
- [Qwen3.5 397B](/models/qwen3-5-397b) - Score: 54.7
- [Qwen3 235B 2507](/models/qwen3-235b-2507) - Score: 53.73
- [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) - Score: 53.39
- [Qwen3.5 Plus](/models/qwen3-5-plus) - Score: 49.89
- [Qwen3.7 Flash](/models/qwen3-7-flash) - Score: 49.01
- [Qwen2.5-1M](/models/qwen2-5-1m) - Score: 47.71
- [Qwen3.6-27B](/models/qwen3-6-27b) - Score: 47.69
- [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) - Score: 43.82
- [Qwen3 Max](/models/qwen3-max) - Score: 40.71
- [Qwen3-Omni-30B-A3B-Instruct](/models/qwen3-omni-30b-a3b-instruct) - Score: 38.33
- [Qwen2.5-VL-32B](/models/qwen2-5-vl-32b) - Score: 37.89
- [Qwen2.5-72B](/models/qwen2-5-72b) - Score: 37.37
- [Qwen2.5 Coder 32B Instruct](/models/qwen2-5-coder-32b-instruct) - Score: 33.37
- [Qwen3.8 Max Preview](/models/qwen3-8-max-preview) - Score: not computed
- [Qwen3-ASR 0.6B](/models/qwen3-asr-0-6b) - Score: not computed
- [Qwen3-ASR 1.7B](/models/qwen3-asr-1-7b) - Score: not computed
- [Qwen-AgentWorld-35B-A3B](/models/qwen-agentworld-35b-a3b) - Score: not computed
- [Qwen3-Omni-30B-A3B-Thinking](/models/qwen3-omni-30b-a3b-thinking) - Score: not computed
- [Qwen2.5-Omni 7B](/models/qwen2-5-omni-7b) - Score: not computed
- [Qwen2-Audio 7B Instruct](/models/qwen2-audio-7b-instruct) - Score: not computed
- [Qwen-Audio 7B](/models/qwen-audio-7b) - Score: not computed
- [Qwen-Audio-Chat 7B](/models/qwen-audio-chat-7b) - Score: not computed
