发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明将蒸馏损失替换为基于相关性的损失,可显著提升语音-音乐编码器在两种合并算法下的合并效果,表明收益源于表示质量。
AI 中文摘要
一个同时处理语音和音乐的紧凑编码器免除了为每个领域维护单独模型的需要。一种实用的方法是将语音教师模型和音乐教师模型蒸馏成两个小型学生模型,然后将它们合并为一个模型。存在两种合并算法:从共享初始化中插值学生权重,或在平均之前对其通道进行排列对齐。先前的工作在这两种算法中固定了蒸馏损失,这留下了损失本身是否影响学生模型合并效果的问题。我们证明它确实有影响。在保持架构和评估不变的情况下,我们将DistilHuBERT损失\Lkd{}替换为基于相关性的损失\Lcl{}。在权重插值下,\Lcl{}学生在18次比较中的17次中更接近其自身更好的端点,并在每个插值权重下在语音任务上领先。在激活排列下,他们在三层需要更少的通道重新排列,其匹配通道相关性更强,并且再次在语音任务上领先。这两种合并算法没有共享机制,但在相同的损失替换下都得到改进。这表明收益在于\Lcl{}产生的表示,而非任何一种算法。
英文摘要
A single compact encoder for both speech and music removes the need to maintain a separate model per domain. A practical recipe distils a speech teacher and a music teacher into two small students, then merges them into one model. Two merging algorithms exist: interpolating the student weights from a shared initialisation, or permuting their channels into alignment before averaging. Prior work fixed the distillation loss in both, leaving open whether the loss itself affects how well the students merge. We show that it does. Holding architecture and evaluation fixed, we replace the DistilHuBERT loss \Lkd{} with a correlation-based loss \Lcl{}. Under weight interpolation, \Lcl{} students stay closer to their own better endpoint in $17$ of $18$ comparisons and lead on the speech tasks at every interpolated weight. Under activation permutation, they need fewer channels rearranged at three layers, their matched channels correlate more strongly, and they again lead on speech. The two merging algorithms share no mechanism, yet both improve under the same loss substitution. This suggests that the gain lies in the representation \Lcl{} produces rather than in either algorithm.
CommentsAccepted at APSIPA ASC 2026