发表机构
DTU Compute; WSA(丹麦技术大学计算系; WSA公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于Mamba扩散模型与扩散薛定谔桥的非成对语音增强框架,无需配对数据,性能与基线相当或更优,推理速度大幅提升。
AI 中文摘要
语音增强(SE)模型通常依赖监督学习,使用成对数据样本,其中干净语音被合成地降质。这种范式在目标环境的特定声学特性未知的真实场景中限制了性能。我们提出了一种完全非成对的SE框架,利用原理性的扩散薛定谔桥(DSB)来学习干净语音分布与降质语音分布之间的随机传输过程。学习传输映射的算法计算量大,因为它们在训练期间通常每一步都需要模拟微分方程。因此,我们提出使用一种专为端到端波形处理设计的高效Mamba扩散模型。我们与最先进的语音增强方法(包括成对和非成对方法)以及经典信号处理算法进行了比较。实验结果表明,我们的方法在性能上与基线持平或更优,同时在推理时快数个数量级。此外,我们展示了DSB公式的灵活性使我们的模型能够跨SE任务泛化,为真实世界的语音恢复提供了稳健且高效的解决方案。
英文摘要
Speech enhancement (SE) models typically rely on supervised learning with paired data examples where clean speech is synthetically degraded. This paradigm limits performance in real-world scenarios where the target environment's specific acoustic characteristics are unknown. We propose a fully unpaired SE framework that uses principled Diffusion Schrödinger Bridges (DSB) to learn a stochastic transport process between a clean and a degraded speech distribution. Algorithms for learning transport maps are computationally heavy since they require simulating differential equations during training, usually at each training step. Therefore, we propose using a high-efficiency Mamba Diffusion Model designed for end-to-end waveform processing. We compare against state-of-the-art methods for speech enhancement, both paired and unpaired, as well as a classical signal processing algorithm. Experimental results show that we are on par or better than the baselines while being orders of magnitude faster during inference. Furthermore, we show that the flexibility of the DSB formulation allows our model to generalize across SE tasks, offering a robust and efficient solution for real-world speech restoration.
Comments15 pages, 4 figures, 4 tables