arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03650eess.AS

面向轻量级声学回声消除中紧凑隐式时频建模的回声感知调制

Echo-Aware Modulation for Compact-Latent Frequency-Time Modeling in Lightweight Acoustic Echo Cancellation

Ye Ni, Ruiyu Liang, Qingyun Wang, Kai Xie, Cairong Zou, Björn W. Schuller

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对轻量级声学回声消除的时频建模局限,提出MSA-EchoLite框架,通过回声感知时频调制模块优化特征表示,以少量额外计算实现更优性能-复杂度权衡,优于现有最优轻量级AEC模型。

中文摘要 AI 辅助

现有的轻量级声学回声消除(AEC)系统常将线性AEC与基于Bark域深度神经网络(DNN)的抑制模块结合,以降低计算开销。这类系统中的下采样层会进一步将输入特征压缩为紧凑瓶颈表示,但这种压缩会削弱时频建模能力并降低性能。为缓解该局限,我们提出MSA-EchoLite,这是一种具有非对称双分支编码器和回声感知时频调制(EAM)模块的轻量级Bark域AEC框架。EAM模块通过对双分支麦克风与回声相关隐式特征间的差异和关联线索进行建模,来丰富压缩后的瓶颈表示。实验结果表明,MSA-EchoLite的Bark域变体相比其频率域变体能实现更优的性能-复杂度权衡,但对特征压缩更敏感。在仅增加非EAM Bark域变体26.1%的浮点运算量(FLOPs)的情况下,经EAM增强的版本达到了频率域变体99.1%的感知语音质量评价(PESQ),而频率域变体需要近两倍的FLOPs,且在信号失真比(SDR)上甚至超越了它。总体而言,MSA-EchoLite仅使用0.2 M参数和100 M FLOPs/s,就优于现有最优的轻量级AEC模型。

英文摘要

Existing lightweight acoustic echo cancellation (AEC) systems often combine linear AEC with Bark-domain DNN-based suppression to lower the computational footprint. In such systems, downsampling layers further compress the input features into a compact bottleneck representation, but this compression weakens frequency-time modeling capacity and degrades performance. To mitigate this limitation, we propose MSA-EchoLite, a lightweight Bark-domain AEC framework with an asymmetric dual-branch encoder and an echo-aware frequency-time modulation (EAM) module. The EAM module enriches the compressed bottleneck representation by modeling discrepancy and correlation cues between the dual-branch microphone and echo-related latent features. Experimental results show that the Bark-domain variant of MSA-EchoLite offers a better performance-complexity trade-off than its frequency-domain counterpart but is more sensitive to feature compression. With only 26.1% additional FLOPs over its non-EAM Bark-domain variant, its EAM-enhanced version achieves 99.1% of the PESQ of the frequency-domain counterpart, which requires nearly twice the FLOPs, and even surpasses it in SDR. Overall, MSA-EchoLite outperforms state-of-the-art lightweight AEC models while using only 0.2 M parameters and 100 M FLOPs/s.

↑