RT-SEMamba:基于渐进式知识蒸馏的实时语音增强Mamba模型
RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation
- Academia Sinica(中央研究院)
- National Taiwan University(台湾大学)
- Kore University of Enna(恩纳科雷大学)
- University of Palermo(巴勒莫大学)
- NVIDIA(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出基于因果时频Mamba模块的全因果语音增强模型RT-SEMamba,通过渐进式知识蒸馏压缩模型,在Voicebank-DEMAND数据集上实现了高质量与低延迟的实时语音增强,速度较教师模型提升2.75倍。
AI中文摘要:
我们提出了RT-SEMamba,这是一种基于因果时频Mamba模块构建的全因果语音增强(SE)模型。与依赖不断增长的键值缓存的Transformer架构不同,Mamba每一层传播固定大小的循环状态,支持内存和带宽高效的长序列推理。我们进一步引入了渐进式知识蒸馏(KD)策略,通过联合蒸馏复杂频谱输出和中间表示,将8层教师模型压缩为浅层1层学生模型。在Voicebank-DEMAND数据集上,8层RT-SEMamba在25 ms算法延迟约束下实现了3.32的PESQ得分,而蒸馏后的1层学生模型在保持相同稳态RTF的情况下,将PESQ得分从朴素1层基线的3.06提升至3.18,同时实现了比教师模型快2.75倍的速度。这些结果表明,结合渐进式KD的状态空间模型可为实时SE提供具有竞争力的质量-延迟权衡。
英文摘要:
We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size recurrent state per layer, enabling memory- and bandwidth-efficient long-form inference. We further introduce a progressive knowledge distillation (KD) strategy that compresses an 8-layer teacher into a shallow 1-layer student by jointly distilling complex spectral outputs and intermediate representations. On Voicebank-DEMAND, the 8-layer RT-SEMamba achieves 3.32 PESQ with a 25 ms algorithmic latency constraint, and the distilled 1-layer student improves over a naive 1-layer baseline from 3.06 to 3.18 PESQ while preserving the same steady-state RTF, delivering a 2.64x speedup over the teacher. These results demonstrate that state-space models with progressive KD provide a competitive quality-latency trade-off for real-time SE.