发表机构
University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究结合TopK稀疏自编码器等方法,在SPEAR和WavLM编码器上实现语言与副语言信息的路由选择性分离,且该分离效果可跨语料库迁移,为语音表示解耦提供了有效方案。
AI 中文摘要
自监督语音编码器在共享的纠缠表示空间中包含语言与副语言信息。我们结合TopK稀疏自编码器、路由特定监督及跨因子对抗器。在冻结的SPEAR和WavLM编码器上,独立探测显示出因子特定的保留与抑制:语言信息在语言路由中保持更强,而副语言因子(包括说话人身份、情感和韵律)在副语言路由中被保留,在语言路由中则大幅减少。在LibriSpeech上学习到的路由组织无需表示侧再训练即可在MSP-Podcast上保持。特征空间路由干预进一步传递了交换的因子,同时在很大程度上保留了未改变路由携带的信息。这些结果在编码器、语料库、独立探测及表示级干预中均显示出一致的路由选择性分离。
英文摘要
Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space. We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries. Across frozen SPEAR and WavLM encoders, independent probes show factor-specific retention and suppression: linguistic information remains stronger in the linguistic route, while paralinguistic factors, including speaker identity, emotion, and prosody, are retained in the paralinguistic route and substantially reduced in the linguistic route. The route organisation learned on LibriSpeech persists on MSP-Podcast without representation-side retraining. Feature-space route interventions further transfer the swapped factor while largely preserving the information carried by the unchanged route. These results show consistent route-selective separation across encoders, corpora, independent probes, and representation-level interventions.
CommentsIn submission