arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

区域感知掩码用于口音鲁棒的跨语言文本到语音合成

Region-Aware Masking for Accent-Robust Cross-Lingual Text-to-Speech

Haoqi Li, Shivam Mehta, Ravi Teja Gadde, Yinghong Lan

arXiv 2610.07524首次发表:更新:

AI 中文总结

针对跨语言TTS口音泄漏,提出区域感知掩码,仅改变掩码位置即可大幅降低口音,且不损可懂度与自然度。

AI 中文摘要

口音泄漏仍然是跨语言零样本文本到语音合成(TTS)中的关键挑战,模型会无意中将说话人的音色和源语言口音同时转移到目标语言输出中。在掩码重建TTS中,这一问题因训练与推理之间的不匹配而被放大:训练时从同语言上下文进行重建,而推理时则基于跨语言上下文进行条件生成。我们通过拼接两种语言的语音片段,并将重建掩码放置在语言边界附近,从而弥合了这一差距,使得训练时呈现出模型在推理时遇到的跨语言提示模式。这无需语言ID、口音标签、平行的同说话人录音或架构更改——仅掩码几何形状不同。在一项由母语者参与的听力测试中,与未适配的基线相比,区域感知掩码将感知口音从0-5量表上的4.3降至0.5;与仅双语适配相比,在已见场景中口音降低18-51%,在未见源语言上降低35-55%,且可懂度或自然度无明显损失。比较掩码几何形状表明,决定口音抑制与说话人保留之间平衡的是掩码的位置而非掩码量:仅覆盖重建区域的掩码对口音抑制最强,但保留参考说话人的可靠性最低,而双区域掩码则能同时实现两者。因此,掩码几何形状是提升口音鲁棒性的一个简单而有效的杠杆。

英文摘要

Accent leakage remains a critical challenge in cross-lingual zero-shot text-to-speech (TTS), where models inadvertently transfer both the speaker's timbre and source-language accent into the target-language output. In mask-reconstruction TTS this is amplified by a train--inference mismatch: training reconstructs from same-language context, while inference conditions on cross-lingual context. We close this gap by concatenating utterances from two languages and placing the reconstruction mask relative to the language boundary, so training presents the cross-lingual prompting pattern the model encounters at inference. This requires no language IDs, accent labels, parallel same-speaker recordings, or architectural changes---only the mask geometry differs. In a listening study with native speakers, region-aware masking reduces perceived accent from 4.3 to 0.5 on a 0-5 scale against the unadapted baseline, and by 18-51\% in seen regimes and 35-55\% on unseen source languages against bilingual adaptation alone, with no measurable loss in intelligibility or naturalness. Comparing mask geometries shows that placement, not the amount masked, governs the balance between accent suppression and speaker preservation: masks confined to the reconstruction region suppress accent most but retain the reference speaker least reliably, whereas two-region masking achieves both. Mask geometry is thus a simple, effective lever for accent robustness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑