发表机构
Nanyang Technological University; Institute for Infocomm Research (I2R)(南洋理工大学; 资讯通信研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过机制性分析揭示,基于相关性的SSL蒸馏中互相关项对角线才是噪声鲁棒性的关键,而自相关项主要改善下游任务,并据此提出蒸馏设计规则。
AI 中文摘要
自监督学习(SSL)语音模型准确但庞大。知识蒸馏对其进行压缩,但学生模型会丧失教师模型的噪声鲁棒性。基于相关性的蒸馏通过两项来解决这一问题:对齐学生与教师表征的互相关项,以及对学生特征进行去相关的自相关项。原始方法和De'HuBERT都将功劳归于自相关项,而未对其进行单独验证。我们证明了相反的结果。在LibriSpeech-100蒸馏中,使用保留的CHiME-3噪声(10 dB),通过无偏的全维度探针、Pearson方差分解、同噪声因果对照以及逐维度分析,识别出互相关对角线是促进噪声不变性的机制,将噪声分类准确率从76.98%降至55.02%,而重新加入自相关项后准确率仍为55.55%。自相关项反而会重组特征空间,提升干净下游任务的准确率,但几乎不消除噪声。在九个语音和音乐任务中,这为蒸馏提供了一个具体的设计规则:对互相关项加权以提升噪声鲁棒性,调整自相关项以获得更好的下游性能,并独立采样教师和学生模型的噪声。
英文摘要
Self-supervised learning (SSL) speech models are accurate but large. Knowledge distillation compresses them, but the student loses the teacher's noise robustness. Correlation-based distillation addresses this with two terms: cross-correlation aligning student and teacher's representations, and self-correlation decorrelating the student's features. Both the original method and De'HuBERT credited the self-correlation term without isolating it. We show the opposite. On LibriSpeech-100 distillation with held-out CHiME-3 noise at 10\,dB, an unbiased full-dimensional probe, a Pearson-variance decomposition, a same-noise causal control, and a per-dimension analysis identify the cross-correlation diagonal as the mechanism that encourages noise invariance, lowering noise-classification accuracy from 76.98\% to 55.02\%, whereas adding the self-correlation term back leaves it at 55.55\%. The self-correlation term instead reorganises the feature space and improves accuracy on the clean downstream tasks, but removes essentially no noise. Across nine speech and music tasks this yields a concrete design rule for distillation: weight the cross-correlation term for noise robustness, tune the self-correlation term for better downstream performance, and sample teacher and student noise independently.
CommentsAccepted at APSIPA ASC 2026