# MAI-Voice-2.1 launch notes: Voice Benchmark Profile

> Microsoft lists two multilingual text-to-speech models at different latency and character-price points.

- Scope: Provider-reported inference latency, languages, character pricing, and a separate Flash end-to-end generation example
- Measurement lanes: Voice experience, Latency
- Primary metric: Model inference latency in milliseconds; USD per million characters
- Available evidence: Specifications published October 1, 2026; no independent listening-test score is included
- Owner: Microsoft AI
- Source snapshots refreshed: 2026-09-30

- Provider report retrieved: 2026-10-01

## Microsoft-reported specifications

Inference timing from the model page; character rates from the October 1 launch.

| Model | Inference latency | Languages | Price per 1M characters |
| --- | --- | --- | --- |
| [MAI-Voice-2.1](/models/mai-voice-2-1) | About 550 ms | 23 | $22 |
| [MAI-Voice-2.1-Flash](/models/mai-voice-2-1-flash) | About 45 ms | 23 | $15 |

## Interpretation limit

The model page reports about 550 ms inference for Voice-2.1 and 45 ms for Flash. The launch gives a separate Flash example: 45 seconds of audio generated with 150 ms end-to-end latency. These timings measure different stages. Microsoft reports 23 languages; its launch says 26 locales while the model page lists 28 labels, so the locale total is not reconciled here. These specifications do not establish an independent speech-quality rank.

## Primary sources

- [Owner page](https://microsoft.ai/models/mai-voice-2-1/)
- [Microsoft audio launch](https://microsoft.ai/news/our-first-streaming-transcription-model/)
- [Azure Speech integration](https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices)

Canonical page: https://benchlm.ai/voice-benchmarks/mai-voice-2-1-launch-notes
