发表机构
KAIST; The University of Tokyo; National Institute of Advanced Industrial Science and Technology(韩国科学技术院; 东京大学; 日本产业技术综合研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过DeMaR模型探究掩码-替换扩散训练在零样本语音合成中的收益,证明其优势源于训练侧噪声增强而非推理时修订,并在LibriTTS上取得更低词错误率。
AI 中文摘要
与自回归模型不同,基于离散扩散的零样本文本到语音模型并行生成语音标记,并能重新审视早期预测。掩码-替换训练通过随机替换部分标记扩展了仅掩码训练,其收益通常归因于自我修正,即修订先前生成标记的能力。然而,训练期间暴露于随机扰动的上下文本身可能改善生成,这引发了一个问题:这些收益是否需要推理时的标记修订。为探究此问题,我们使用DeMaR,它结合了掩码-替换训练与置信度排序的仅掩码采样,同时保持总训练破坏概率不变。在LibriTTS上从头训练,DeMaR在相同语音分词器下取得了比自回归和仅掩码扩散基线更低的词错误率(WER)。当每个标记在首次被解除掩码后保持不变时,这一优势依然存在。在两种异构语音分词器上匹配训练条件表明,在限制条件下,噪声上下文增强和替换监督均改善了WER。这些发现证明了替换在训练侧的益处,超越了仅用于启用推理时标记修订的作用。
英文摘要
Unlike autoregressive models, discrete diffusion-based models for zero-shot text-to-speech generate speech tokens in parallel and can revisit earlier predictions. Mask-and-replace training extends mask-only training by randomly replacing some tokens, and its gains are commonly attributed to self-correction, the ability to revise previously generated tokens. However, exposure to randomly perturbed context during training may itself improve generation, raising the question of whether these gains require inference-time token revision. To investigate this question, we use DeMaR, which combines mask-and-replace training with confidence-ranked mask-only sampling while preserving the total training corruption probability. Trained from scratch on LibriTTS, DeMaR achieves lower word error rates (WER) than autoregressive and mask-only diffusion baselines using the same speech tokenizer. This advantage persists when each token remains unchanged after first being unmasked. Matched training conditions on two heterogeneous speech tokenizers show that both noisy-context augmentation and replacement supervision improve WER under this restriction. These findings demonstrate training-side benefits of replacement beyond enabling inference-time token revision.
Comments5 pages, 1 figure, 3 tables. Under review