CoFi-Lite:突破超轻量级语音增强的极限
CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement
浏览论文内容
中文总结 AI 辅助
研究旨在突破超轻量级语音增强模型复杂度极限,提出CoFi-Lite模型,通过解耦频谱建模、利用并行编解码路径及跨路径融合模块,实现低计算资源需求,性能优于基线且降低计算成本。
中文摘要 AI 辅助
超轻量级模型对基于深度学习的语音增强算法在边缘设备上的部署至关重要。尽管近期方法在计算复杂度和性能间取得一定平衡,但进一步突破复杂度极限需更精细设计。本文提出CoFi-Lite,一种高效模型,将频谱建模解耦为粗粒度和细粒度流。利用两个并行对称编解码路径,同时提取全带包络和低频细节进行互补增强。还引入新型跨路径融合模块促进特征交互。CoFi-Lite计算资源需求极低,实验结果表明其性能优于超轻量级基线GTCRN,计算复杂度仅为其40.26%,放大版本性能与SOTA超轻量级模型AdaptCRN相当且计算成本降低19.34%。
英文摘要
Ultra-lightweight models are essential for the deployment of deep learning-based speech enhancement algorithms on edge devices. Although recent approaches have achieved a certain balance between computational complexity and performance, pushing the complexity limits further demands more sophisticated designs. In this letter, we propose CoFi-Lite, a highly efficient model that decouples spectral modeling into coarse- and fine-grained streams. By leveraging two parallel and symmetric encoder-decoder paths, it simultaneously extracts full-band envelopes and low-frequency details for complementary enhancement. In addition, a novel Cross-Path Fusion (CPF) module is introduced to bridge the distinct paths, facilitating efficient feature interaction. Remarkably, CoFi-Lite requires extremely low computational resources, featuring only 12.87M MACs/s and 83.12k parameters. Experimental results demonstrate that our proposed model outperforms the ultra-lightweight baseline GTCRN while requiring only 40.26% of its computational complexity. Its scaled-up variant also delivers performance on par with that of the SOTA ultra-lightweight model AdaptCRN alongside a 19.34% reduction in computational cost. Audio examples are available at https://acceleration123.github.io/CoFiLite-demo/.