# MAI-Voice-2.1-Flash Benchmark Scores & Performance

> MAI-Voice-2.1-Flash has MAI-Voice-2.1 launch notes results. Voice evidence is display-only and does not produce a text-model score.

Canonical page: https://benchlm.ai/models/mai-voice-2-1-flash

Last updated: October 1, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Microsoft |
| Source Type | Proprietary |
| Reasoning Type | Non-Reasoning |
| Context Window | N/A |
| Official model card | [Microsoft model card](https://microsoft.ai/pdf/MAI-Voice-2.1-Flash-Model-Card-Memo.pdf) |
| Overall Score | Not scored (voice evidence only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: MAI-Voice-2.1
- Variant: flash
- Benchmarks covered: 0 of 645
- Sibling models: [MAI-Voice-2.1](/models/mai-voice-2-1)
- Coverage note: Speech and voice results appear below; no weighted text-model benchmark rows are stored.

## Speech and voice evidence

[MAI-Voice-2.1 launch notes](/voice-benchmarks/mai-voice-2-1-launch-notes)

The model page reports about 550 ms inference for Voice-2.1 and 45 ms for Flash. The launch gives a separate Flash example: 45 seconds of audio generated with 150 ms end-to-end latency. These timings measure different stages. Microsoft reports 23 languages; its launch says 26 locales while the model page lists 28 labels, so the locale total is not reconciled here. These specifications do not establish an independent speech-quality rank.

| Metric | Result |
| --- | --- |
| Languages | 23 |
| Model inference latency | About 45 ms |
| Price | $15 per 1M characters |

Microsoft-reported specifications, October 1, 2026; inference is separate from end-to-end generation.

- [Evaluation source](https://microsoft.ai/models/mai-voice-2-1/)

Released October 1, 2026. Azure Speech documents public-preview access through SSML voice selectors and the Speech SDK, with Voice Live support. The launch also lists OpenRouter, Microsoft Foundry, MAI Playground, and Vercel; LiveKit is coming soon. Both versions support 23 languages and gated voice cloning.

## Other Microsoft Models

- [Phi-4](/models/phi-4) - Score: 26.37
- [MAI-Thinking-1](/models/mai-thinking-1) - Score: not computed
- [MAI-Code-1.1-Flash](/models/mai-code-1-1-flash) - Score: not computed
- [Raw Phi-4 mini direct logits](/models/raw-phi-4-mini) - Score: not computed
- [Fara1.5-27B](/models/fara-1-5-27b) - Score: not computed
- [Fara1.5-4B](/models/fara-1-5-4b) - Score: not computed
- [MAI-Transcribe-2-Streaming](/models/mai-transcribe-2-streaming) - Score: not computed
- [MAI-Voice-2.1](/models/mai-voice-2-1) - Score: not computed
- [MAI-Transcribe-2](/models/mai-transcribe-2) - Score: not computed
- [VibeVoice-ASR-Streaming 1.5B](/models/vibevoice-asr-streaming-1-5b) - Score: not computed
- [VibeVoice-ASR-Streaming 7B](/models/vibevoice-asr-streaming-7b) - Score: not computed
- [MAI-Cyber-1-Flash](/models/mai-cyber-1-flash) - Score: not computed
- [MAI-Voice-2-Flash](/models/mai-voice-2-flash) - Score: not computed
- [MAI-Transcribe-1.5](/models/mai-transcribe-1-5) - Score: not computed
- [MAI-Voice-2](/models/mai-voice-2) - Score: not computed
- [Phi-4 Multimodal Instruct](/models/phi-4-multimodal-instruct) - Score: not computed
