Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Official provider results

Qwen3.8-Omni-Flash omni evaluation

Qwen’s September 18, 2026 launch post reports a 32-row audio and audio-visual table for Qwen3.8-Omni-Flash against Qwen3.5-Omni-Plus, Gemini 3.8 Flash, Seed 2.0 Lite, and Muse Spark 1.2, plus a separate static-versus-agent comparison on three video sets.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
Audio-visual agent benchmarks (WildClawBench-MM, UniClawBench, AgenticVBench, OmniGAIA), audio-visual understanding, reasoning, captioning and interaction sets (DailyOmni, WorldSense, AVUT, JoinAVBench, OmniVideoBench, Video-MME-v2, LVOmniBench, OmniCloze, OmniCap-IF, QIVD, StreamingBench), and audio sets covering multi-speaker ASR, multilingual ASR and speech translation, audio understanding and grounding, music understanding, and spoken interaction
Primary metric
Accuracy or task score (higher is better), except the ASR rows, which report DER, cpWER, and WER (lower is better)
Owner
Qwen Team
Available evidence
Exact figures tabulated in the official launch post for all five compared systems

What it measures

Spoken reasoning
Does the system understand and answer spoken requests?
Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Voice experience
Is the exchange natural, robust, and responsive?

Available results

WildClawBench-MM
71.0
#1 of 4 reported

Multimodal tool use with the Claude Code harness, over the multimodal subset of WildClawBench. Qwen3.5-Omni-Plus 34.5, Gemini 3.8 Flash 58.9, Seed 2.0 Lite 41.9.

UniClawBench
69.6
#1 of 4 reported

Multimodal tool use with the OpenClaw harness. Qwen3.5-Omni-Plus 67.1, Gemini 3.8 Flash 69.0, Seed 2.0 Lite 61.2.

AgenticVBench
36.8
#2 of 4 reported

Multimodal tool use with the Claude Code harness. Gemini 3.8 Flash leads at 45.0; Qwen3.5-Omni-Plus 14.5, Seed 2.0 Lite 10.0.

OmniGAIA
74.0
#2 of 4 reported

Web search with no agent harness. Gemini 3.8 Flash leads at 78.6; Seed 2.0 Lite 64.4, Qwen3.5-Omni-Plus 57.2.

OmniVideoBench
63.4
#3 of 5

Audio-visual reasoning, static understanding. Gemini 3.8 Flash 65.2, Muse Spark 1.2 62.2, Seed 2.0 Lite 58.5, Qwen3.5-Omni-Plus 53.8. With the Qwen Code agent harness the score rises to 67.8 while tokens per query fall from 145,736 to 79,117.

Video-MME-v2
65.0
#3 of 4 reported

Audio-visual reasoning, static understanding. Gemini 3.8 Flash 71.0, Seed 2.0 Lite 64.9, Qwen3.5-Omni-Plus 47.9. With the Qwen Code agent harness the score rises to 71.3.

LVOmniBench
63.3
#2 of 3 reported

Long video reasoning, static understanding. Gemini 3.8 Flash leads at 70.7; Qwen3.5-Omni-Plus 53.2. With the Qwen Code agent harness the score rises to 73.6, ahead of Gemini 3.8 Flash at 70.7.

DailyOmni
85.1
Best-equal

Audio-visual understanding. Tied with Qwen3.5-Omni-Plus; Gemini 3.8 Flash 84.0, Seed 2.0 Lite 81.4, Muse Spark 1.2 79.6.

JoinAVBench
75.9
#1 of 5

Audio-visual understanding. Qwen3.5-Omni-Plus 74.1, Muse Spark 1.2 71.8, Seed 2.0 Lite 70.6, Gemini 3.8 Flash 70.4.

WorldSense
68.5
#2 of 5

Audio-visual understanding. Gemini 3.8 Flash leads at 69.6; Seed 2.0 Lite 67.3, Muse Spark 1.2 65.0, Qwen3.5-Omni-Plus 63.9.

AVUT
86.6
#2 of 5

Audio-visual understanding. Gemini 3.8 Flash leads at 88.0; Qwen3.5-Omni-Plus 85.9, Muse Spark 1.2 82.4, Seed 2.0 Lite 81.5.

OmniCap-IF
CSR 80.6 / ISR 28.2
#2 of 5 on both

Audio-visual captioning. Gemini 3.8 Flash leads on both at CSR 81.9 / ISR 28.3; Muse Spark 1.2 77.9 / 26.8, Seed 2.0 Lite 74.6 / 18.1, Qwen3.5-Omni-Plus 72.1 / 14.1.

OmniCloze
63.2
#3 of 5

Audio-visual captioning. Muse Spark 1.2 leads at 65.3; Qwen3.5-Omni-Plus 64.2, Gemini 3.8 Flash 60.9, Seed 2.0 Lite 56.3.

QIVD
69.6
#1 of 5

Audio-visual interaction. Gemini 3.8 Flash 69.1, Qwen3.5-Omni-Plus 65.6, Seed 2.0 Lite and Muse Spark 1.2 62.0.

StreamingBench
80.8
#1 of 5

Audio-visual interaction. Gemini 3.8 Flash 79.9, Muse Spark 1.2 77.8, Seed 2.0 Lite 77.2, Qwen3.5-Omni-Plus 57.1.

AliMeeting Test DER / cpWER
3.4 / 17.2
#1 of 5 on both

Multi-speaker ASR, lower is better. Seed 2.0 Lite 75.1 / 76.1, Gemini 3.8 Flash 72.6 / 53.1, Qwen3.5-Omni-Plus 88.1 / 89.6, Muse Spark 1.2 93.7 / 92.7.

AISHELL-4 DER / cpWER
2.8 / 11.2
#1 of 5 on both

Multi-speaker ASR, lower is better. Seed 2.0 Lite 64.8 / 64.2, Gemini 3.8 Flash 66.4 / 56.9, Muse Spark 1.2 91.3 / 86.0, Qwen3.5-Omni-Plus 100.0 / 100.0.

MagicData-RAMC DER / cpWER
5.7 / 14.1
#1 of 5 on both

Multi-speaker ASR, lower is better. Seed 2.0 Lite 43.4 / 35.1, Gemini 3.8 Flash 67.9 / 33.8, Muse Spark 1.2 82.1 / 75.3, Qwen3.5-Omni-Plus 98.4 / 97.1.

MLC-SLM (en) DER / cpWER
4.0 / 14.2
#1 of 5 on both

Multi-speaker ASR, lower is better. Seed 2.0 Lite 40.4 / 45.5, Gemini 3.8 Flash 60.8 / 26.6, Qwen3.5-Omni-Plus 68.6 / 63.9, Muse Spark 1.2 74.3 / 52.9.

WenetSpeech (Net) WER
4.8
#3 of 5

ASR, lower is better. Qwen3.5-Omni-Plus leads at 3.7; Seed 2.0 Lite 4.3, Gemini 3.8 Flash 14.2, Muse Spark 1.2 68.2.

WenetSpeech (Meeting) WER
4.6
#1 of 5

ASR, lower is better. Seed 2.0 Lite 4.7, Qwen3.5-Omni-Plus 4.8, Gemini 3.8 Flash 16.7, Muse Spark 1.2 42.6.

FLEURS-ASR WER
9.3
#3 of 5

Multilingual ASR over 60 languages, lower is better. Qwen3.5-Omni-Plus leads at 7.2; Gemini 3.8 Flash 7.9, Muse Spark 1.2 23.6, Seed 2.0 Lite 32.1.

FLEURS-S2TT BLEU
31.8
#3 of 5

Multilingual speech translation over 60 languages. Gemini 3.8 Flash leads at 33.0; Qwen3.5-Omni-Plus 32.2, Muse Spark 1.2 28.8, Seed 2.0 Lite 24.8.

SpotSoundBench
67.2
#1 of 5

Audio grounding. Qwen3.5-Omni-Plus 64.2, Seed 2.0 Lite 59.6, Gemini 3.8 Flash 39.7, Muse Spark 1.2 16.9.

MMAU
81.8
#2 of 5

Audio understanding. Qwen3.5-Omni-Plus leads at 81.9; Seed 2.0 Lite 77.2, Gemini 3.8 Flash 76.9, Muse Spark 1.2 63.5.

MMAR
79.8
Best-equal

Audio understanding. Tied with Qwen3.5-Omni-Plus; Gemini 3.8 Flash 78.5, Seed 2.0 Lite 77.7, Muse Spark 1.2 67.3.

MMSU
82.1
#3 of 5

Audio understanding. Gemini 3.8 Flash leads at 83.3; Qwen3.5-Omni-Plus 83.0, Seed 2.0 Lite 80.2, Muse Spark 1.2 59.9.

LongAudioSpan accuracy / rubric / chain
82.7 / 71.8 / 48.2
#1 of 3 reported on accuracy and rubric

Long audio reasoning. Gemini 3.8 Flash 79.3 / 65.5 / 64.6 leads the chain metric; Qwen3.5-Omni-Plus 74.4 / 49.8 / 45.1.

MuchoMusic-RUL
72.6
#1 of 5

Music understanding. Qwen3.5-Omni-Plus 71.6, Seed 2.0 Lite 61.7, Gemini 3.8 Flash 53.7, Muse Spark 1.2 40.1.

HumMusQA
75.8
#1 of 5

Music understanding. Qwen3.5-Omni-Plus 75.5, Gemini 3.8 Flash 71.2, Seed 2.0 Lite 66.0, Muse Spark 1.2 63.3.

MusTBench
50.6
#1 of 5

Music understanding. Qwen3.5-Omni-Plus 49.1, Seed 2.0 Lite 44.0, Gemini 3.8 Flash 40.3, Muse Spark 1.2 29.4.

Audio MultiChallenge
71.5
#2 of 5

Spoken interaction. Gemini 3.8 Flash leads at 71.9; Seed 2.0 Lite 63.4, Muse Spark 1.2 57.9, Qwen3.5-Omni-Plus 57.6.

WildSpeech
74.3
#4 of 5

Spoken interaction. Gemini 3.8 Flash leads at 76.4; Qwen3.5-Omni-Plus 75.7, Seed 2.0 Lite 74.5, Muse Spark 1.2 73.4.

VoiceBench
91.6
#3 of 5

Spoken interaction. Qwen3.5-Omni-Plus leads at 92.9; Gemini 3.8 Flash 92.3, Seed 2.0 Lite 84.1, Muse Spark 1.2 79.8.

Interpretation limit

Qwen ran every system in this table itself and chose the harness for each agent row: WildClawBench-MM and AgenticVBench use Claude Code, UniClawBench uses OpenClaw, OmniGAIA uses no harness, and the static-versus-agent comparison uses Qwen Code. WildClawBench-MM covers only the multimodal subset of WildClawBench. Competitor runs use provider-specific media settings (Gemini 3.8 Flash at media_resolution=high, Seed 2.0 Lite at max_frame_tokens=384), so the comparison is not independently reproducible. These voice-protocol results stay separate from BenchLM’s weighted text-model ranking.