多样本合成监督用于口音转换
Multi-sample Synthetic Supervision for Accent Conversion
浏览论文内容
中文总结 AI 辅助
提出多样本合成监督框架,通过联合评估候选波形构建口音转换目标,提升口音分类准确率2.49个百分点,保持说话人相似度,实现7.81%词错误率。
中文摘要 AI 辅助
口音转换(AC)需要在保留说话人身份和语言内容的同时改变口音,然而平行录音数据稀缺。语音合成为监督提供了替代来源,但生成的目标在口音实现和源保留方面存在差异。我们提出了一种多样本合成监督框架,该框架通过联合评估每个源-口音条件下的候选波形中的这些属性来构建转换目标。选定的候选目标提供关联的离散语音编码,代表语言内容、韵律和说话风格,作为口音条件自回归自适应的目标。与每个训练示例使用一个生成目标相比,我们的方法将目标口音分类准确率提高了2.49个百分点,而说话人相似度几乎保持不变,词错误率增加了0.29个百分点。在六种目标口音中,我们的方法实现了7.81%的词错误率,并在评估系统中获得了最高的平均听力评分。
英文摘要
Accent conversion (AC) requires changing accent while preserving speaker identity and linguistic content, yet parallel recordings are scarce. Speech synthesis provides an alternative source of supervision, but generated targets vary in accent realization and source preservation. We propose a multi-sample synthetic supervision framework that constructs conversion targets by jointly assessing these properties across candidate waveforms for each source--accent condition. Selected candidates provide associated discrete speech codes, representing linguistic content, prosody, and speaking style, as targets for accent-conditioned autoregressive adaptation. Compared with using one generated target per training example, our method improves target-accent classification accuracy by 2.49 percentage points, while speaker similarity remains nearly unchanged, and word error rate increases by 0.29 percentage points. Across six target accents, our method achieves 7.81% WER and the highest mean listening ratings among evaluated systems.