Dispatch
Speech recognition models reproduce known transcript errors and hallucinate silenced numbers, suggesting benchmark scores overstate real ability
Hugging Face measured benchmark optimization across 11 ASR systems including Whisper-Large-V3, Parakeet-TDT, Qwen3-ASR-0.6B and Cohere-Transcribe-03-2026. Six of 11 reproduced an erroneous VoxPopuli transcript that omitted 'Thank you' despite the audio containing it, with the lowest-WER models reproducing benchmark errors at the highest rates (18-30%). Silencing numbers in LibriSpeech audio still saw top models 'recover' them 30-40% of the time, and models matched dataset-specific orthography ('Mr.' vs 'Mister') at up to 90% accuracy against a 50% baseline.
↳ Follow the thread