Official provider results
VibeVoice-ASR-Streaming model notes
Microsoft’s model card documents streaming speaker-attributed transcription, hotword support, and ten-language coverage; evaluation results are published as an image.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- Capabilities and language coverage from the model card
- Primary metric
- Documented capabilities (no scored benchmark)
- Owner
- Microsoft Research
- Available evidence
- Evaluation table published only as an image
What it measures
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Available results
- Transcription
- Streaming, speaker-attributed
- Languages
- 10
Documented
Who said what as speech arrives.
Documented
Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish.
Interpretation limit
No numeric results are transcribed because the card publishes them as an image.