Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Official provider results

Canto launch evaluation

Wispr’s September 17, 2026 launch post reports word error rates for Canto against five competing transcription systems on two private dictation sets and three public English datasets.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
10 hours of opt-in Wispr Flow dictations from 2,300+ speakers, a 3-hour audio challenge set (noisy, low-volume, and short-utterance audio), and the English subsets of FLEURS, LibriSpeech, and Common Voice
Primary metric
Word error rate (lower is better)
Owner
Wispr Advanced Interfaces Lab
Available evidence
Exact figures labelled on the charts in the official launch post; the two dictation sets are private to Wispr

What it measures

Voice experience
Is the exchange natural, robust, and responsive?

Available results

Real-world dictation WER
3.4%
#1 of 6

10 hours of Wispr Flow dictations. The rest of Wispr’s comparison: Gemini 3.1 Pro 3.6%, GPT Transcribe 4.0%, AssemblyAI Universal-3.5 Pro 4.2%, Gemini 3.5 Transcribe 4.5%, Deepgram Nova-3 7.4%.

Audio challenge set WER
9.1%
#2 of 6

Gemini 3.1 Pro leads at 8.2%; GPT Transcribe 11.3%, Gemini 3.5 Transcribe 12.2%, AssemblyAI Universal-3.5 Pro 13.9%, Deepgram Nova-3 23.5%. Canto is the best of the real-time models Wispr tested.

Noisy-audio WER
7.6%
#2 of 6

Low-SNR dictations with traffic, wind, or competing speech. Gemini 3.1 Pro leads at 6.1%; GPT Transcribe 8.9%, Gemini 3.5 Transcribe 10.0%, AssemblyAI 10.4%, Deepgram Nova-3 22.2%.

Low-volume WER
10.1%
Best-equal

Whispered and far-field speech. Tied with Gemini 3.1 Pro; GPT Transcribe 13.2%, Gemini 3.5 Transcribe 14.1%, AssemblyAI 18.3%, Deepgram Nova-3 25.0%.

Short-dictation WER
21.4%
Best-equal

One- and two-word utterances, the hardest subset for every system. Tied with Gemini 3.1 Pro and AssemblyAI; Gemini 3.5 Transcribe 25.0%, GPT Transcribe 33.6%, Deepgram Nova-3 41.4%.

LibriSpeech WER
1.7%
Best-equal

Tied with AssemblyAI; Gemini 3.1 Pro and GPT Transcribe 1.8%, Gemini 3.5 Transcribe 2.1%, Deepgram Nova-3 2.9%.

FLEURS WER
3.6%
#5 of 6

AssemblyAI leads at 2.9%; GPT Transcribe 3.2%, Gemini 3.1 Pro and Gemini 3.5 Transcribe 3.5%, Deepgram Nova-3 8.6%.

Common Voice WER
9.2%
#3 of 6

AssemblyAI leads at 8.7% and Gemini 3.1 Pro follows at 8.8%; Gemini 3.5 Transcribe and GPT Transcribe 10.7%, Deepgram Nova-3 26.0%.

Interpretation limit

Wispr ran every system in this comparison itself, and the two headline sets are its own unreleased dictation data, so the results are not independently reproducible. Competitor runs used no contextual prompting, which is not how several of them are deployed, and the challenge-set winner (Gemini 3.1 Pro) is a frontier multimodal model Wispr describes as unsuitable for real-time dictation.