arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22824eess.AS

自适应深度与专家细化用于高效语音增强

Adaptive Depth and Expert Refinement for Efficient Speech Enhancement

Xikun Lu, Yujian Ma, Yunda Chen, Xianquan Jiang, Jinqiu Sang

首次发表
浏览论文内容

中文总结 AI 辅助

针对固定深度神经语音增强系统计算冗余的问题,提出ADER框架,通过自适应深度控制器和条件专家路由器实现输入依赖的细化,在VCTK-DEMAND上以更少参数和计算达到高语音质量。

中文摘要 AI 辅助

大多数神经语音增强系统对所有输入使用固定的处理深度,当较少的细化步骤足够时,这可能会引入不必要的计算。我们提出了自适应深度与专家细化(ADER),一种具有输入依赖计算的参数共享渐进增强框架。ADER结合了用于硬性提前终止的自适应深度控制器(ADC)和条件专家路由器(CER),后者在每次执行的细化迭代中选择一个轻量级残差适配器。我们进一步引入了退出感知的中间监督(EIS),以直接优化用于提前退出的候选中间输出。在VCTK-DEMAND上,ADER将MP-SENet的参数数量和平均计算量分别减少了70.4%和51.3%,同时实现了3.37的WB-PESQ。总体而言,ADER实现了输入依赖的细化,并在推理期间减少了冗余计算。

英文摘要

Most neural speech enhancement systems use a fixed processing depth for all inputs, which can introduce unnecessary computation when fewer refinement steps are sufficient. We propose Adaptive Depth and Expert Refinement (ADER), a parameter-shared progressive enhancement framework with input-dependent computation. ADER combines an Adaptive Depth Controller (ADC) for hard early termination with a Conditional Expert Router (CER) that selects one lightweight residual adapter at each executed refinement iteration. We further introduce Exit-aware Intermediate Supervision (EIS) to directly optimize candidate intermediate outputs for early exit. On VCTK-DEMAND, ADER reduces the parameter count and average computation of MP-SENet by 70.4% and 51.3%, respectively, while achieving a WB-PESQ of 3.37. Overall, ADER enables input-dependent refinement and reduces redundant computation during inference.

补充信息

↑