Triplets-like Russian Figure Skaters: Can Kullback-Leibler Divergence be Used Tell Their Difference?
Speaker adaptation is one of those ASR problems that sounds solved until you actually try to adapt a sequence-to-sequence model without wrecking the encoder's hard-won acoustic representations.
Speaker adaptation is one of those ASR problems that sounds solved until you actually try to adapt a sequence-to-sequence model without wrecking the encoder's hard-won acoustic representations. This talk uses the memorable frame of near-identical Russian figure skaters to ask whether KL divergence can distinguish very similar speaker distributions, and walks through "Listen, Attend, Spell and Adapt": a speaker-adapted seq2seq ASR architecture that keeps the base model intact while learning a lightweight per-speaker component.
Expect a careful look at how KL regularization prevents catastrophic drift during adaptation, how adaptation data budgets trade off against WER gains, and why LAS-style attention encoders behave differently under adaptation than hybrid HMM systems. For anyone building personalized voice interfaces, voice cloning stacks, or dictation products with per-user acoustic profiles, the ideas here still translate directly to Conformer and transducer-based systems. Hit play if speaker adaptation is more than a curiosity in your ASR stack.
