arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

校正强迫:生成式语音增强中扩散模型与流模型的统一后训练

Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement

Qing Yao, Lijian Gao, Qirong Mao

arXiv 2609.24651首次发表:更新:

发表机构

Jiangsu University; Jiangsu Engineering Research Center of Big Data Ubiquitous Perception and Intelligent Agriculture Applications; Provincial Key Laboratory of Computational Intelligence and New Technologies in Low-Altitude Digital Agriculture(江苏大学; 江苏省大数据泛在感知与智能农业应用工程技术研究中心; 省低空数字农业计算智能与新技术重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散与流模型在语音增强中训练-推理不匹配问题,提出校正强迫后训练范式,通过自生成滚动校正预测,提升感知质量与重建保真度。

AI 中文摘要

扩散模型和流模型作为语音增强中有前景的生成范式,面临训练与推理不匹配的问题:训练使用解析路径状态,而推理则沿离散化采样轨迹对自生成的滚动状态递归评估模型。这种不匹配导致预测误差和离散化误差累积。为解决此问题,我们引入校正强迫(CoF),一种后训练范式,迫使扩散模型和流模型从自生成的滚动中学习并校正其预测。CoF在动态采样调度下将滚动状态上的干净语音预测向真实值校正,使模型暴露于不同的推理条件。它进一步利用局部校正的反事实转移作为事实转移的参考来正则化局部演化。通过共享的干净语音预测参数化表达模型输出,CoF将相同的后训练目标应用于扩散和流公式。使用SB-VE和OT-CFM的实验表明,感知质量和重建保真度均有提升,且在不同采样步数下表现稳健。

英文摘要

Diffusion and flow models, as promising generative paradigms for speech enhancement, face a training--inference mismatch: training uses analytical path states, whereas inference recursively evaluates models on self-generated rollout states along discretized sampling trajectories. This mismatch causes prediction and discretization errors to accumulate. To address it, we introduce Corrective Forcing (CoF), a post-training paradigm that forces diffusion and flow models to learn from self-generated rollouts and correct their predictions. CoF corrects clean-speech predictions on rollout states toward the ground truth under dynamic sampling schedules, exposing the model to varying inference conditions. It further regularizes local evolution using locally corrected counterfactual transitions as references for factual transitions. By expressing model outputs through a shared clean-speech prediction parameterization, CoF applies the same post-training objective across diffusion and flow formulations. Experiments with SB-VE and OT-CFM demonstrate improvements in perceptual quality and reconstruction fidelity, together with robust performance across different numbers of sampling steps.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑