Official provider results
MAI-Voice-2.1 launch notes
Microsoft lists two multilingual text-to-speech models at different latency and character-price points.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- Provider-reported inference latency, languages, character pricing, and a separate Flash end-to-end generation example
- Primary metric
- Model inference latency in milliseconds; USD per million characters
- Owner
- Microsoft AI
- Available evidence
- Specifications published October 1, 2026; no independent listening-test score is included
What it measures
Voice experience
Is the exchange natural, robust, and responsive?
Latency
How long do responses, tool calls, and full tasks take?
Available results
Microsoft-reported specifications
Inference timing from the model page; character rates from the October 1 launch.
| Model | Inference latency | Languages | Price per 1M characters |
|---|---|---|---|
| MAI-Voice-2.1 | About 550 ms | 23 | $22 |
| MAI-Voice-2.1-Flash | About 45 ms | 23 | $15 |
Interpretation limit
The model page reports about 550 ms inference for Voice-2.1 and 45 ms for Flash. The launch gives a separate Flash example: 45 seconds of audio generated with 150 ms end-to-end latency. These timings measure different stages. Microsoft reports 23 languages; its launch says 26 locales while the model page lists 28 labels, so the locale total is not reconciled here. These specifications do not establish an independent speech-quality rank.