Skip to main content
BenchLM
Official provider results

MAI-Voice-2.1 launch notes

Microsoft lists two multilingual text-to-speech models at different latency and character-price points.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
Provider-reported inference latency, languages, character pricing, and a separate Flash end-to-end generation example
Primary metric
Model inference latency in milliseconds; USD per million characters
Owner
Microsoft AI
Available evidence
Specifications published October 1, 2026; no independent listening-test score is included

What it measures

Voice experience
Is the exchange natural, robust, and responsive?
Latency
How long do responses, tool calls, and full tasks take?

Available results

Microsoft-reported specifications

Inference timing from the model page; character rates from the October 1 launch.

Inference timing from the model page; character rates from the October 1 launch.
ModelInference latencyLanguagesPrice per 1M characters
MAI-Voice-2.1About 550 ms23$22
MAI-Voice-2.1-FlashAbout 45 ms23$15

Interpretation limit

The model page reports about 550 ms inference for Voice-2.1 and 45 ms for Flash. The launch gives a separate Flash example: 45 seconds of audio generated with 150 ms end-to-end latency. These timings measure different stages. Microsoft reports 23 languages; its launch says 26 locales while the model page lists 28 labels, so the locale total is not reconciled here. These specifications do not establish an independent speech-quality rank.