用于文本引导语音修复与编辑的多编解码器离散扩散模型
Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing
- Ben-Gurion University of the Negev(内盖夫本-古里安大学)
- University of Haifa(海法大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出基于分层编解码器令牌的离散扩散框架SIEDD,其核心架构HiCoDD结合音素级约束等技术,在RealEdit基准中实现最优语音编辑性能,且优于自回归基线,显著提升上下文保留的语音重建与编辑效果。
AI中文摘要:
语音录音常包含缺失、损坏或错误的区域,需在不重新合成整个话语的情况下对其进行重建或修改。语音修复用于恢复缺失片段,而语音编辑则根据编辑后的文本替换口语内容。这两项任务都要求生成的语音表达出预期的词汇,同时与周围的说话人身份、韵律、时间和录音条件保持一致。离散扩散特别适合这些任务,因为它可以在同时受左右两侧声学上下文约束的情况下迭代优化掩码令牌。我们提出SIEDD,这是一个用于文本引导语音修复与编辑的离散扩散框架,基于分层编解码器令牌构建。其核心架构HiCoDD遵循RVQ生成顺序,将先前生成的码本表示为干净、已提交的声学上下文,仅对当前优化码本应用扩散。这种分离实现了无泄漏的联合训练,同时匹配顺序的由粗到精推理。该模型还结合了音素级约束、跨度定位的无分类器引导以及时长预测,以支持固定时长修复和可变时长文本编辑。在RealEdit基准测试中,SIEDD在评估方法中取得了最佳的整体语音编辑性能,并且在所有语音修复设置(包括单间隙和多间隙)中,均优于评估的自回归基线方法。这些结果表明,显式建模编解码器层次结构可显著提升保留上下文的语音重建与编辑效果。完整代码可访问此https URL。
英文摘要:
Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing replaces spoken content according to an edited transcript. Both tasks require the generated speech to express the intended words while remaining consistent with the surrounding speaker identity, prosody, timing, and recording conditions. Discrete diffusion is particularly well suited to these tasks because it can iteratively refine masked tokens while jointly conditioning on both left and right acoustic context. We introduce SIEDD, a discrete diffusion framework for text-guided speech inpainting and editing over hierarchical codec tokens. Its core architecture, HiCoDD, follows the RVQ generation order by representing previously generated codebooks as clean, committed acoustic context and applying diffusion only to the current refinement codebook. This separation enables leakage-free joint training while matching sequential coarse-to-fine inference. The model further combines phoneme-level conditioning, span-localized classifier-free guidance, and duration prediction to support both fixed-duration inpainting and variable-duration text edits. On the RealEdit benchmark, SIEDD achieves the best overall speech-editing performance among the evaluated methods. It also outperforms the evaluated autoregressive baselines across all speech-inpainting settings, on both single and multiple gaps. These results demonstrate that explicitly modeling the codec hierarchy substantially improves context-preserving speech reconstruction and editing. See our full code at https://github.com/iftachShoham/SIEDD.