# Qwen3.8-Omni-Flash omni evaluation: Voice Benchmark Profile

> Qwen’s September 18, 2026 launch post reports a 32-row audio and audio-visual table for Qwen3.8-Omni-Flash against Qwen3.5-Omni-Plus, Gemini 3.8 Flash, Seed 2.0 Lite, and Muse Spark 1.2, plus a separate static-versus-agent comparison on three video sets.

- Scope: Audio-visual agent benchmarks (WildClawBench-MM, UniClawBench, AgenticVBench, OmniGAIA), audio-visual understanding, reasoning, captioning and interaction sets (DailyOmni, WorldSense, AVUT, JoinAVBench, OmniVideoBench, Video-MME-v2, LVOmniBench, OmniCloze, OmniCap-IF, QIVD, StreamingBench), and audio sets covering multi-speaker ASR, multilingual ASR and speech translation, audio understanding and grounding, music understanding, and spoken interaction
- Measurement lanes: Spoken reasoning, Task completion, Conversation dynamics, Voice experience
- Primary metric: Accuracy or task score (higher is better), except the ASR rows, which report DER, cpWER, and WER (lower is better)
- Available evidence: Exact figures tabulated in the official launch post for all five compared systems
- Owner: Qwen Team
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

Qwen ran every system in this table itself and chose the harness for each agent row: WildClawBench-MM and AgenticVBench use Claude Code, UniClawBench uses OpenClaw, OmniGAIA uses no harness, and the static-versus-agent comparison uses Qwen Code. WildClawBench-MM covers only the multimodal subset of WildClawBench. Competitor runs use provider-specific media settings (Gemini 3.8 Flash at media_resolution=high, Seed 2.0 Lite at max_frame_tokens=384), so the comparison is not independently reproducible. These voice-protocol results stay separate from BenchLM’s weighted text-model ranking.

## Primary sources

- [Code](https://github.com/QwenLM/Qwen-MM-Plugins)
- [Owner page](https://qwen.ai/blog?id=qwen3.8-omni-flash)

Canonical page: https://benchlm.ai/voice-benchmarks/qwen3-8-omni-flash-evaluation
