AI 中文总结
本研究证明在预训练中增强语言判别能力可缩小多语言语音模型与单语言模型的性能差距,同时保留跨语言共享,支持语言判别对降低多语言学习成本的因果作用。
AI 中文摘要
多语言自监督语音模型可以通过跨语言信息共享获益,但在匹配的总预训练数据预算下,其性能仍不及单语言模型。我们证明,在预训练期间增强模型区分语言的能力,可以减少并在某些指标上消除连续音素和更高级别语言度量上的多语言差距,同时保留大量的跨语言共享。在受控的英语/法语HuBERT设置中,我们测试了两种增强语言判别的干预措施:辅助语言分类器和每种语言的k-means目标。在各项干预措施中,连续特征音素判别错误(phone-ABX,越低越好)从双语基线的11.6%降至10.4%(单语言:10.8%),而词汇性能(sWUGGY,越高越好)从52.1%提升至56.7%(单语言:58.5%),韵律性能(ProsAudit,词汇子任务,越高越好)从68.9%提升至72.9%(单语言:72.6%)。在HuBERT训练阶段中,大多数语言度量上的最大提升发生在第一轮迭代引入语言判别时,而较晚或重复的干预措施带来的改进较小,并伴随着语言间分离度的增加。这些结果支持语言判别在减少多语言学习额外成本中的因果作用。
英文摘要
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing. Using a controlled English/French HuBERT setting, we test two interventions which strengthen language discrimination: an auxiliary language classifier and per-language k-means targets. Across interventions, continuous-feature phone discrimination error (phone-ABX, lower is better) decreases from 11.6% in the bilingual baseline to 10.4% (monolingual: 10.8%), while lexical performance (sWUGGY, higher is better) increases from 52.1% to 56.7% (monolingual: 58.5%) and prosodic performance (ProsAudit, lexical subtask, higher is better) from 68.9% to 72.9% (monolingual: 72.6%). Across HuBERT training stages, the strongest gains on most linguistic measures occur when language discrimination is introduced in the first iteration, whereas later or repeated interventions yield smaller improvements and are accompanied by increased language-wise segregation. These results support a causal role for language discrimination in reducing the additional cost of multilingual learning.