发表机构
Technische Universität Braunschweig; GN Advanced Science(布伦瑞克工业大学; GN前沿科学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出低延迟因果扩散模型DiffVQE2/DiffVQE2-S,用于联合声学回声和噪声控制,在ICASSP 2023 AEC挑战赛盲测集上超越最先进判别模型,且复杂度更低,支持流式处理。
AI 中文摘要
免提通信设备和扬声器电话固有地受到声学回声和背景噪声的影响。为了减轻这些干扰,端到端判别式训练的神经网络已成为研究和部署中表现最佳的方法。尽管生成方法的最新进展已为各种语音增强任务提供了显著成果,但基于扩散的声学回声控制(AEC)研究仍局限于非因果、话语级处理,因此在实践中应用不广。通过这项工作,我们首次提出了低延迟(即因果)的基于扩散的联合AEC和噪声控制模型DiffVQE2 / DiffVQE2-S,在多个客观指标上,尤其是在主观MOS方面,分别超越了目前最先进的DeepVQE / DeepVQE-S模型。此外,我们的模型复杂度更低。进一步地,我们展示了将有限的lookahead应用于高效的DiffVQE2-S模型可以实现更高的性能。这些结果是在ICASSP 2023 AEC挑战赛盲测集上获得的。我们声称这是首个支持流式处理的、基于扩散的声学回声和噪声控制方法,其性能超越了最先进的判别式方法。
英文摘要
Hands-free communication devices and speakerphones are inherently affected by acoustic echo and background noise. To mitigate these impairments, end-to-end discriminatively trained neural networks have emerged as the best-performing approach in research and deployment. While recent advancements in generative methods have provided remarkable results for various speech enhancement tasks, diffusion-based acoustic echo control (AEC) research is still restricted to non-causal, utterance-level processing, thereby not widely applicable in practice. With this work, we are the first to propose low-delay (i.e., causal) diffusion-based joint AEC and noise control models DiffVQE2 / DiffVQE2-S, excelling the so-far state of the art DeepVQE / DeepVQE-S models in multiple objective metrics, and, most importantly, in subjective MOS, respectively. In addition, our models are less complex. Furthermore, we show that a limited lookahead applied to the efficient DiffVQE2-S model allows for an even higher performance. These results have been obtained on the ICASSP 2023 AEC Challenge blind test set. We claim the first streaming-capable, diffusion-based acoustic echo and noise control that excels state-of-the-art discriminative approaches.
Commentsaccepted at IWAENC 2026