RAWD-TTS:离散扩散语音克隆的无比例奖励对齐
RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning
浏览论文内容
中文总结 AI 辅助
针对离散扩散语音克隆中监督训练与波形级属性不匹配的问题,提出RAWD-TTS方法,利用识别和说话人奖励及组相对优势加权掩码重建,无需目标音频,在500个俄语提示上显著降低词错误率并提升说话人相似度。
中文摘要 AI 辅助
零样本文本到语音合成能够从短参考录音中生成说话人声音的新话语。语音克隆需要准确的内容和保留的说话人身份,但监督式声学标记预测并不直接优化这些波形级属性。基于奖励的后训练解决了这一不匹配问题,但在离散扩散中,标记选择和揭示位置共同定义了采样轨迹,使对齐复杂化。我们引入了RAWD-TTS(无比例优势加权去噪),它使用识别和说话人奖励对解码样本进行评分,并利用组相对优势对这些样本的掩码标记重建进行加权,无需反向轨迹似然或目标音频。在500个俄语CV3-Eval语音克隆提示上,联合对齐在奖励选择检查点将词错误率从3.18%降至2.42%(相对降低24.0%),在最终检查点降至2.58%(相对降低19.0%),而WavLM说话人余弦相似度从0.733提升至0.748和0.755。受控实验表征了识别-身份权衡以及损坏数量、组组成和加权的影响。
英文摘要
Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.
发表机构
- lab260
- BitmanagerAI
- MTUCI
机构由 AI 辅助整理,请以论文原文为准。