Official provider results
Granite Speech 5.0 TurboCTC model notes
IBM’s model card and Granite 4.2 blog document the model’s size, training data, and throughput; Open ASR leaderboard results are published as images.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- Parameter count, training hours, and RTFx throughput on a single H200
- Primary metric
- Throughput (higher is better)
- Owner
- IBM
- Available evidence
- Throughput published in text; WER charts are images
What it measures
Latency
How long do responses, tool calls, and full tasks take?
Available results
- Parameters
- 470M
- Throughput
- ~12,600 RTFx on one H200
- Training data
- ~60,000 hours English
Documented
Conformer encoder with block self-attention; no LLM backbone.
Provider-reported
About three hours of audio per second.
Documented
Public corpora including CommonVoice, MLS, and YODAS.
Interpretation limit
IBM publishes the Open ASR and FFASR leaderboard comparisons only as images, so no WER value is transcribed here.