arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12099cs.SDcs.CL

RT-SEMamba:基于渐进式知识蒸馏的实时语音增强Mamba模型

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

  • Academia Sinica(中央研究院)
  • National Taiwan University(台湾大学)
  • Kore University of Enna(恩纳科雷大学)
  • University of Palermo(巴勒莫大学)
  • NVIDIA(英伟达公司)

机构由 AI 辅助整理,请以论文原文为准。

Rong Chao, Sung-Feng Huang, Moreno La Quatra, Sabato Marco Siniscalchi, Wen-Huang Cheng, Szu-Wei Fu, Yu Tsao

AI总结:

该研究提出基于因果时频Mamba模块的全因果语音增强模型RT-SEMamba,通过渐进式知识蒸馏压缩模型,在Voicebank-DEMAND数据集上实现了高质量与低延迟的实时语音增强,速度较教师模型提升2.75倍。

AI中文摘要:

我们提出了RT-SEMamba,这是一种基于因果时频Mamba模块构建的全因果语音增强(SE)模型。与依赖不断增长的键值缓存的Transformer架构不同,Mamba每一层传播固定大小的循环状态,支持内存和带宽高效的长序列推理。我们进一步引入了渐进式知识蒸馏(KD)策略,通过联合蒸馏复杂频谱输出和中间表示,将8层教师模型压缩为浅层1层学生模型。在Voicebank-DEMAND数据集上,8层RT-SEMamba在25 ms算法延迟约束下实现了3.32的PESQ得分,而蒸馏后的1层学生模型在保持相同稳态RTF的情况下,将PESQ得分从朴素1层基线的3.06提升至3.18,同时实现了比教师模型快2.75倍的速度。这些结果表明,结合渐进式KD的状态空间模型可为实时SE提供具有竞争力的质量-延迟权衡。

英文摘要:

We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size recurrent state per layer, enabling memory- and bandwidth-efficient long-form inference. We further introduce a progressive knowledge distillation (KD) strategy that compresses an 8-layer teacher into a shallow 1-layer student by jointly distilling complex spectral outputs and intermediate representations. On Voicebank-DEMAND, the 8-layer RT-SEMamba achieves 3.32 PESQ with a 25 ms algorithmic latency constraint, and the distilled 1-layer student improves over a naive 1-layer baseline from 3.06 to 3.18 PESQ while preserving the same steady-state RTF, delivering a 2.64x speedup over the teacher. These results demonstrate that state-space models with progressive KD provide a competitive quality-latency trade-off for real-time SE.

补充信息

↑