← Machine Learning Street Talk (MLST)

Speech Recognition Is Not a Solved Problem — Pavan Kumar Reddy

Machine Learning Street Talk (MLST)2026年9月15日1時間42分

Speech Recognition Is Not a Solved Problem — Pavan Kumar Reddy

Machine Learning Street Talk (MLST)

0:001:42:22
このエピソードはアーカイブのため、日本語要約の対象外です。
番組の概要欄(原文)

<p>Pavan Kumar Reddy leads audio research at Mistral AI. He joins Tim Scarfe for a deep technical tour of Voxtral — and explains why the frontier of deployed voice is still a cascade of specialised models rather than one end-to-end system.</p><p><br></p><p>IN PARTNERSHIP WITH MISTRAL AI:</p><p>---</p><p>This episode was produced in partnership with Mistral AI.</p><p>Mistral AI: https://mistral.ai/</p><p>---</p><p><br></p><p>The conversation opens on architecture. Voxtral Chat feeds a 3B Ministral text trunk with continuous embeddings from an audio encoder, passed to the decoder as direct token input rather than through cross-attention as in Whisper, so the model can answer questions about emotion, timing and who spoke when without an intermediate transcript to lose them. The real-time model becomes a dual-stream decoder that consumes audio and emits text at once, at a target delay down to 160ms, with slower streams in parallel for anything that can wait for more context.</p><p><br></p><p>On generation, Pavan explains why Voxtral TTS predicts continuous latents rather than discrete codec tokens, traces the lineage from SoundStream through EnCodec to Mimi&#39;s split of semantic and acoustic codebooks, and places FSQ and flow matching in it. Tim presses on the priors underneath: why a mel spectrogram instead of raw waveform, what noise augmentation buys, and when acoustic overfitting becomes somebody&#39;s fine-tuning problem. Then the failure modes. Diarisation is emitted autoregressively inside the transcript rather than by a separate head, which makes streaming diarisation fragile — less context, late speaker changes, invented extra speakers. And because the architecture commits to what it has already predicted, one out-of-distribution mistake compounds into looping or skipped segments, which is what DPO corrects: the negative supervision pre-training and SFT cannot give.</p><p><br></p><p>The last third is the argument Tim keeps returning to. Customers running voice agents over millions of sessions describe scaffolding, not a solved problem, with a sharp drop outside the top few languages. Cascades survive because each component stays separately adaptable, observable and constrainable. And voice alone is cognitive debt: absorbing information and deciding in one serial stream is harder than glancing at a menu. Voice becomes ubiquitous beside a screen, not instead of one.</p><p><br></p><p>---</p><p>TIMESTAMPS:</p><p>00:00:00 Cold open</p><p>00:00:46 Why Mistral moved into audio</p><p>00:09:27 Inside Voxtral: trunk, encoder, dual streams</p><p>00:20:22 Speech that works in real time</p><p>00:30:52 How a voice becomes tokens</p><p>00:39:59 Flow matching, FSQ and the new codec</p><p>00:52:51 When speech models lose the speaker</p><p>01:03:23 Correcting hallucinations with preferences</p><p>01:12:12 Controlling synthetic speech</p><p>01:20:06 Why cascades still win</p><p>01:29:25 Speech in the wild</p><p>01:33:46 Audio models as interfaces</p><p>01:37:54 Why voice still needs a screen</p><p><br></p><p>---</p><p>REFERENCES:</p><p>paper:</p><p>[00:01:42] Mistral 7B</p><p>https://arxiv.org/abs/2310.06825</p><p>[00:09:38] Voxtral</p><p>https://arxiv.org/abs/2507.13264</p><p>[00:14:41] Whisper: Robust Speech Recognition</p><p>https://arxiv.org/abs/2212.04356</p><p>[00:19:11] Voxtral Realtime</p><p>https://arxiv.org/abs/2602.11298</p><p>[00:21:52] Delayed Streams Modeling (Kyutai)</p><p>https://arxiv.org/abs/2509.08753</p><p>[00:30:52] Voxtral TTS</p><p>https://arxiv.org/abs/2603.25551</p><p>[00:32:38] SoundStream neural audio codec</p><p>https://arxiv.org/abs/2107.03312</p><p>[00:34:59] Flow Matching for Generative Modeling</p><p>https://arxiv.org/abs/2210.02747</p><p>[00:37:03] EnCodec: High Fidelity Neural Audio Compression</p><p>https://arxiv.org/abs/2210.13438</p><p>[00:37:42] Moshi and the Mimi codec</p><p>https://arxiv.org/abs/2410.00037</p><p>[00:39:05] Finite Scalar Quantization (FSQ)</p><p>https://arxiv.org/abs/2309.15505</p><p>[01:03:33] Direct Preference Optimization (DPO)</p><p>https://arxiv.org/abs/2305.18290</p><p>dataset:</p><p>[00:46:14] Mozilla Common Voice</p><p>https://commonvoice.mozilla.org/en/datasets</p><p>organization:</p><p>[00:50:47] Hugging Face</p><p>https://huggingface.co/</p>

X でシェアSpotify で聴くApple Podcasts で聴く

関連エピソード