arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多语言语音模型中的音系干扰

Phonological Interference in Multilingual Speech Models

Moran Yanuka, Raja Giryes, Moris Alper

arXiv 2610.11275首次发表:更新:

发表机构

Tel Aviv University; University of Miami(特拉维夫大学; 迈阿密大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究发现多语言语音模型存在音系干扰问题,提出窗口化语言估计(WLE)修复方法,可消除代码转换输入中34%-69%的干扰且不影响单语言性能。

AI 中文摘要

音素级模型将语音转录或生成为音素序列,音素是区分单词的最小声音单位。这些模型支持细粒度的发音控制与理解,但在输入不匹配任何单一训练语言时(如两种语言交替的代码转换语音,或训练中未包含的低资源语言)常出现故障。我们识别出其背后的系统性故障模式:音系干扰,即模型假设输入属于单一语言并施加该语言的音系,覆盖与假设语言冲突的局部音素级决策。我们通过模型保留某语言独有、另一语言缺失的音素的频率来衡量干扰。在代码转换输入上,两款音素识别器(语音转音素模型)和一款音素条件文本转语音模型会丢失32%至79%的这类音素,但丢失两种语言共有音素的比例低得多。在未见过的语言上,我们发现音素识别器会施加其分配给该语音的训练语言的音系,且分配置信度越高,丢失未见过语言独有、分配语言缺失的音素越多。我们从模型内部激活中探测其语言估计,并将干扰追溯至低维子空间。在单语言语音上,将该子空间转向另一语言会使模型丢失仅原语言使用的音素,生成仅目标语言有的音素。我们引入窗口化语言估计(WLE),这是一种推理时修复方法,用每个位置周围短窗口计算的语言估计替换该子空间中的模型语言估计。在代码转换输入上,WLE消除了所有三款模型中34%至69%的干扰,且在识别器中基本不改变单语言性能。

英文摘要

Phoneme-level models transcribe or generate speech as a sequence of phonemes, the smallest sound units that distinguish words. These models enable fine-grained pronunciation control and understanding, yet often fail on input that does not match any single training language, such as speech alternating between two languages, known as code-switching, or low-resource languages absent from training. We identify a systematic failure mode behind this, phonological interference: models assume the input is in a single language and impose its phonology, overriding local phoneme-level decisions that conflict with the assumed language. We measure interference by how often a model retains phonemes that one language has but the other lacks. On code-switched input, two phone recognizers (speech-to-phoneme models) and a phoneme-conditioned text-to-speech model lose 32% to 79% of these phonemes, but lose far fewer of the phonemes both languages share. On unseen languages, we find that phone recognizers impose the phonology of the training language they assign to the speech, and the more confident the assignment, the more they lose phonemes the unseen language has but the assigned language lacks. We probe the models' language estimate from their internal activations, and trace interference to a low dimensional subspace. On monolingual speech, steering this subspace toward another language makes the model lose the phonemes that only the original language uses and produce phonemes that only the target language has. We introduce windowed language estimation (WLE), an inference time repair that replaces the model's language estimate in this subspace with one computed from a short window around each position. On code-switched input, WLE removes 34% to 69% of the interference in all three models, and in the recognizers it leaves monolingual performance essentially unchanged.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑