口音类比引导:跨语言语音克隆中同等口音下更高的说话人相似度
Accent Analogy Guidance: More Speaker Similarity at Equal Accent in Cross-Lingual Voice Cloning
浏览论文内容
中文总结 AI 辅助
针对跨语言语音克隆中的口音泄漏问题,提出训练无关的口音类比引导(AAG)方法,通过减去模型自身预测估计的口音方向,在保持口音不变的同时提升说话人相似度,并在多个开源TTS模型上验证了有效性。
中文摘要 AI 辅助
在跨语言零样本文本到语音合成中,参考语音的口音会渗入目标语音。我们提出口音类比引导(AAG),这是一种无需训练的采样器项,它从模型自身对同一合成声音在两种语言下的预测中估计口音方向并减去该方向,从而抵消声音成分,仅保留口音。通过使用真实配音数据上的盲法LLM口音评判器,对参考语音和文本之间的分类器自由引导进行重新加权及其变体,它们都保持在同一条身份-口音权衡曲线附近;我们根据方法在同等口音下高于该曲线的说话人相似度(ΔSIM)进行评分。在四个开源TTS模型上,AAG均位于曲线上方:在OmniVoice上,三个测试集的ΔSIM为+0.11至+0.27(口音评分在1-5分制上为3.51至4.28,说话人相似度为0.29,而重新加权仅保持0.02);MaskGCT和CosyVoice 2也位于其曲线上方,在F5-TTS上,AAG比任何重新加权设置都更接近母语。无LLM的语言识别度量和十二人听者小组的结果一致。前提测试和模型自身曲线的可达范围可提前指示AAG能否获益及大致获益程度,并预测了唯一无法获益的模型(X-Voice)。
英文摘要
In cross-lingual zero-shot text-to-speech, the accent of the reference leaks into the target speech. We propose accent analogy guidance (AAG), a training-free sampler term that subtracts an accent direction estimated from the model's own predictions for one synthetic voice rendered in both languages, so the voice cancels and only the accent remains. By a blind LLM accent judge on real dubbing data, reweighting classifier-free guidance between reference and text, and its variants, stay near one identity-accent trade-off curve; we score a method by its speaker similarity above that curve at equal accent ($Δ$SIM). Across four open TTS models AAG lies above the curve: on OmniVoice $Δ$SIM is +0.11 to +0.27 on three test sets (accent 3.51 to 4.28 on a 1-5 scale at speaker similarity 0.29, where reweighting keeps 0.02); MaskGCT and CosyVoice 2 also lie above their curves, and on F5-TTS it is more native than any reweighting setting. An LLM-free language-ID measure and a twelve-listener panel agree. A premise test and the reach of a model's own curve indicate in advance whether and roughly how much AAG can gain, predicting the one model where it gains nothing (X-Voice).
发表机构
- ESTsoft
机构由 AI 辅助整理,请以论文原文为准。