DriftSE:基于生成漂移的语音增强
DriftSE: Speech Enhancement with Generative Drifting
浏览论文内容
中文总结 AI 辅助
DriftSE提出一种基于生成漂移的单步语音增强框架,通过双潜在并行漂移同时保持语音可懂度和声学保真度,实现非配对训练,并在四个数据集上以1 NFE达到最先进的词错误率。
中文摘要 AI 辅助
我们提出了DriftSE,一种新颖的单步生成式语音增强框架,其被形式化为一个潜在分布均衡问题。在训练期间,漂移场通过在潜在域中的漂移,将生成器的推前分布与干净语音流形对齐。在推理期间,漂移过程被丢弃,从而实现单步生成。我们确定其增强质量从根本上取决于潜在表示的选择。语义潜在表示保留了语音结构,但无法捕捉物理声学线索,而声学潜在表示重建了物理信号,但存在语言幻觉的风险。因此,我们引入了双潜在漂移,在语义和声学潜在表示中并行进行漂移,以同时保留语音可懂度和声学保真度。此外,我们证明了DriftSE通过对齐潜在分布而非精确的逐点目标,实现了完全非配对训练。因此,在缺乏配对的带噪-干净样本的情况下,DriftSE促进了跨数据集学习。此外,DriftSE在多种生成器骨干网络上展现出广泛的架构灵活性。在加性去噪和卷积去混响上的广泛评估表明,在离线和实时因果设置下均具有稳健的单步增强性能。值得注意的是,DriftSE在严格以1 NFE运行时,在所有四个评估数据集上均达到了最先进的词错误率。代码和音频示例可在线获取。
英文摘要
We propose DriftSE, a novel one-step generative framework for speech enhancement formulated as a latent distribution equilibrium problem. During training, the drifting field aligns the generator's pushforward distribution with the clean speech manifold through drifting in a latent domain. During inference, the drifting process is discarded, enabling one-step generation. We establish that its enhancement quality depends fundamentally on the choice of latent representation. Semantic latents preserve phonetic structure but fail to capture physical acoustic cues, whereas acoustic latents reconstruct the physical signal but risk linguistic hallucination. Therefore, we introduce dual-latent drifting, performing parallel drifting in both semantic and acoustic latents to simultaneously preserve phonetic intelligibility and acoustic fidelity. Additionally, we demonstrate that DriftSE enables fully unpaired training by aligning latent distributions rather than exact point-wise targets. Consequently, DriftSE facilitates cross-dataset learning in the absence of paired noisy-clean samples. Moreover, DriftSE exhibits broad architectural flexibility across different generator backbones. Extensive evaluations on additive denoising and convolutive dereverberation demonstrate robust one-step enhancement across both offline and real-time causal settings. Notably, DriftSE achieves state-of-the-art word error rates across all four evaluated datasets while strictly operating at 1 NFE. Code and audio examples are available online.
发表机构
- Victoria University of Wellington(惠灵顿维多利亚大学)
- GN Advanced Science(GN 前沿科学)
- Lincoln University(林肯大学)
机构由 AI 辅助整理,请以论文原文为准。