speech recognition
4 published articles
AI Research
Voice Memory: a 776-byte file that tells speech recognition when to do nothing
A new inference-only scheme for speech recognition learns restraint: a frozen corrector reads a per-domain memory file and decides when to abstain. Unconstrained correction breaks correct tokens on up to 64% of edits; Voice Memory cuts that to 35% and lowers weighted WER from 8.36% to 7.52%.
2026-08-06
Artificial Intelligence
480ms latency, Whisper-level accuracy, an open recipe, Nvidia's Canary redefines real-time ASR
Nvidia's Nemo Canary matches Whisper-level accuracy on LibriSpeech and Common Voice while streaming at 480ms. The open source release includes the full training recipe, making it a rare reproducible contender in real-time ASR.
2026-07-26
Open Source AI
480 ms and open source: a streaming model that finally matches Whisper's quality
Voxtral Realtime matches offline transcription quality at sub-second latency and is open source. The 13-language model uses a novel causal audio encoder and is trained end-to-end for streaming rather than adapted from offline systems.
2026-07-17
Artificial Intelligence
Nvidia's new audio model does five jobs at once and beats the specialists at their own game
Nvidia's Audex unifies audio understanding, generation, and text reasoning in a single model, matching or beating task-specific systems on speech and audio benchmarks without sacrificing text performance.
2026-07-09