缓解自监督语音表征中的口音-语言混淆以实现语言识别
Mitigating Accent-Language Confusion in Self-Supervised Speech Representations for Language Identification
浏览论文内容
中文总结 AI 辅助
针对自监督语音表征中口音与语言混淆问题,提出基于几何投影的L1偏差去除方法,仅用母语语音估计偏差方向,在五个MMS-LID模型上显著提升L2口音语音识别,且无需L2数据或模型适配。
中文摘要 AI 辅助
口语语言识别(LID)旨在识别目标语言,而不受口音影响。然而,在实践中,从自监督语音表征微调而来的LID模型经常将口音与语言混淆,将非母语(L2)语音错误分类为说话者的第一语言(L1)。我们表明,非母语语音表征位于母语目标语言和母语L1极点之间,导致系统性错误分类。为解决这一问题,我们引入了一种几何投影方法,该方法仅从母语语音中估计L1偏差方向,并在冻结的LID头部之前将其移除。在五个MMS-LID模型和非母语语料库上,这种投影显著提高了对L2口音语音的目标语言识别,同时保持了对母语语音的预测。这些结果表明,口音引起的L1偏差可以在表征空间内直接校正,无需L2训练数据或模型适配。
英文摘要
Spoken language identification (LID) aims to recognize the target language regardless of accent. In practice, however, LID models fine-tuned from self-supervised speech representations frequently confuse accents with languages, misclassifying non-native (L2) speech as the speaker's first language (L1). We show that non-native speech representations lie between native target-language and native L1 poles, causing systematic misclassification. To address this, we introduce a geometric projection that estimates an L1-bias direction solely from native speech and removes it before the frozen LID head. Across five MMS-LID models and non-native corpora, this projection substantially improves target language identification for L2-accented speech while preserving predictions for native speech. These results show that accent-induced L1 bias can be corrected directly within the representation space without L2 training data or model adaptation.
发表机构
- University of Southern California(南加州大学)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。