# Voxtral Mini 4B Realtime 2602 Benchmark Scores & Performance

> Voxtral Mini 4B Realtime 2602 has Fermion speech evaluation results. Voice evidence is display-only and does not produce a text-model score.

Canonical page: https://benchlm.ai/models/voxtral-mini-4b-realtime-2602

Last updated: September 30, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Mistral AI |
| Source Type | Open Weight |
| Reasoning Type | Non-Reasoning |
| Context Window | N/A |
| Official model card | [Mistral AI model documentation](https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602) |
| Overall Score | Not scored (voice evidence only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Voxtral Mini 4B Realtime 2602
- Variant: base
- Benchmarks covered: 0 of 486
- Coverage note: Speech and voice results appear below; no weighted text-model benchmark rows are stored.

## Speech and voice evidence

[Phonon-2 launch evaluation](/voice-benchmarks/phonon-2-launch-evaluation)

Fermion ran Phonon-2, Phonon-1, and Parakeet Redux with the Open ASR Leaderboard code on full test sets. The other accuracy rows are leaderboard figures quoted by Fermion, without independent re-verification here. Some model cards report different runs; those values are not mixed into this comparison. Voxtral's download size is an estimate at 16 bits. Throughput includes log-mel processing and excludes model load; batch-128 GPU results are separate from single-stream results. The Mac runtime comparison uses 20 dictations totaling 797 seconds at each runtime's defaults. Download size does not measure runtime memory. These results do not enter text-model or voice-agent rankings.

| Metric | Result |
| --- | --- |
| Seven-set mean WER | 6.12% |
| LibriSpeech clean WER | 1.62% |
| LibriSpeech other WER | 4.94% |
| AMI WER | 13.34% |
| Earnings-22 WER | 9.31% |
| GigaSpeech WER | 8.80% |
| SPGISpeech WER | 2.23% |
| VoxPopuli WER | 2.60% |
| Download | About 8,000 MB (estimated) |

Leaderboard row quoted by Fermion; Fermion launch table retrieved September 30, 2026. Lower WER is better; throughput excludes model load.

- [Evaluation source](https://www.fermionresearch.com/research/phonon-2/)

Open streaming speech transcription weights under Apache 2.0. The official model card documents 13 languages, configurable transcription delay, BF16 weights, and vLLM inference. Fermion estimates the download size from the parameter count at 16 bits.

## Other Mistral AI Models

- [Voxtral 4B TTS 2603](/models/voxtral-4b-tts-2603) - Score: not computed
