arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21188eess.AS

SlimDiffuSE:使用可瘦身网络实现高效的基于扩散模型的语音增强

SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks

Nagashree K. S. Rao, Shrishti Saha Shetu, Mohamed Elminshawi, Emanuël A. P. Habets, Andreas Brendel

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出SlimDiffuSE可瘦身扩散模型,采用贪心搜索算法优化网络宽度调度,在PESQ、SI-SDR指标无显著下降的情况下,将计算复杂度最高降低87.5%,实现高效语音增强。

中文摘要 AI 辅助

基于扩散的模型在语音增强领域逐渐兴起,并在多个基准数据集上达到了最先进的性能。扩散模型的一个主要缺点是数据生成需要对通常规模较大的神经网络进行多次评估,这导致整体计算复杂度很高。在本研究中,我们提出了一种可瘦身扩散模型,该模型在整个数据生成过程中采用自适应网络宽度以降低计算成本。通过使用贪心搜索算法优化网络宽度调度,我们的方法实现了与基准扩散模型相当的性能,同时计算复杂度显著降低。值得注意的是,我们的方法将计算复杂度降低了高达87.5%,且在感知语音质量评估(PESQ)和SI-SDR等客观指标上没有显著下降。

英文摘要

Diffusion-based models are emerging in the speech enhancement domain and are achieving state-of-the-art performance across various benchmark datasets. A major downside of diffusion models is that data generation requires many evaluations of a typically large neural network, which results in high overall complexity. In this work, we propose a slimmable diffusion model that employs adaptive network widths throughout the data generation process to reduce computational cost. By using a greedy search algorithm to optimize the network width schedule, our method achieves performance comparable to baseline diffusion models with significantly reduced computational complexity. Notably, our approach reduces the computational complexity by up to $87.5\%$ without a significant drop in objective metrics, such as perceptual evaluation of speech quality (PESQ) and SI-SDR.

↑