Hacker News
Wispr's Canto Is an ASR Model Trained on Messy Real Dictation Rather Than Clean Corpora
Wispr Advanced Interfaces Lab announced Canto on September 17, built and evaluated against real Wispr Flow dictations recorded in offices, commutes and meetings across varied microphones and background noise rather than studio audio. On 10 hours of real dictation it posted the lowest word error rate against Google, OpenAI, AssemblyAI and Deepgram, and on a 3-hour challenge set it led among real-time models, with Gemini 3.1 Pro scoring better overall but not usable in real time. On public benchmarks it only ties for lowest on LibriSpeech and does not lead FLEURS or Common Voice, which is the honest admission the post makes: the model is tuned for the distribution it ships into.
↳ Follow the thread