arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.26350cs.SDcs.CL

基于神经音频编解码器令牌剖析自监督语音学习对训练语言的敏感性

Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

  • National Institute of Advanced Industrial Science and Technology (AIST)(日本国立先进工业科学技术研究所)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell, William Chen, Satoru Fukayama, Shinji Watanabe

AI总结:

该研究通过控制变量分析发现,基于神经音频编解码器的自监督语音模型的下游性能依赖SSL预训练语言,不依赖编解码器训练语言,据此提出单个编解码器可跨语言复用,需对齐预训练语言与目标语言。

AI中文摘要:

神经音频编解码器(NACs)因能将语音表示为离散令牌而备受关注,除压缩外,这些离散令牌还可用于训练自监督学习(SSL)模型,这类基于编解码器的SSL模型降低了数据存储与计算成本,实现了可扩展的SSL预训练。然而,其语言敏感性尚不明确,当语言发生变化时,基于编解码器的SSL模型可能需要重新训练,这损害了其效率。本文通过固定其中一项,改变NAC训练语言或SSL预训练语言,对语言敏感性展开系统分析。实验结果表明,下游任务性能对NAC训练语言不敏感,但强烈依赖于SSL预训练语言。这些发现说明,单个NAC可跨语言复用,而使SSL预训练语言与目标语言对齐至关重要。

英文摘要:

Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be used to train self-supervised learning (SSL) models. Such models, referred to as codec-based SSL models, reduce data storage and computational cost, enabling scalable SSL pre-training. However, their language sensitivity remains unclear. When the language changes, codec-based SSL models may require retraining, which undermines their efficiency. In this paper, we present a systematic analysis of language sensitivity by varying either the NAC training language or the SSL pre-training language while keeping the other fixed. Experimental results show that downstream performance is insensitive to the NAC training language but strongly dependent on the SSL pre-training language. These findings suggest that a single NAC can be reused across languages, while aligning the SSL pre-training language with the target language is crucial.

补充信息

↑