arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23038cs.SD

LipsAM:用于收敛即插即用音频信号恢复的利普希茨连续神经网络

LipsAM: Lipschitz-continuous Neural Networks for Convergent Plug-and-Play Audio Signal Recovery

Kazuki Matsumoto, Ren Uchida, Natsuki Yoshino, Kohei Yatabe

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有理论框架无法适配音频处理常用DNN的问题,提出利普希茨连续的LipsAM架构,开发其利普希茨常数评估框架,并将其应用于即插即用音频恢复,实现了可收敛的语音去混响算法。

中文摘要 AI 辅助

深度神经网络(DNN)的利普希茨连续性是为其行为建立理论保证的关键。从理论和实践角度出发,已提出多种方法来构建利普希茨连续架构并控制其利普希茨常数。然而,音频信号处理中常用的多种DNN架构超出了现有理论框架的适用范围,阻碍了声学应用中利普希茨连续模型的发展。特别是,尽管处理复值信号幅度和相位的DNN被广泛采用,但在现有框架下无法实现利普希茨连续性。本文为解决这一局限,为构建幅度修正器(AM,一类仅对复值输入幅度进行操作的DNN架构)建立了具有可证明利普希茨连续性的理论基础。具体而言,本文推导了AM为利普希茨连续的充要条件,并提出了对应于音频信号常用架构(包括时频掩码)的LipsAM(利普希茨连续AM)。此外,本文开发了评估其利普希茨常数的高效框架,并针对部分提出的架构解析推导了这些常数。作为应用,本文提出了用于即插即用(PnP)音频信号恢复的CoReM-LipsAM(基于LipsAM的受控残差映射),将DNN作为数据驱动先验集成到基于模型的信号处理算法中。所得PnP算法的收敛性由CoReM-LipsAM架构在结构上保证,并通过语音去混响实验进行了经验验证。

英文摘要

The Lipschitz continuity of deep neural networks (DNNs) is essential for establishing theoretical guarantees regarding their behavior. From both theoretical and practical perspectives, various methods have been proposed to construct Lipschitz-continuous architectures and control their Lipschitz constants. However, several DNN architectures common in audio signal processing fall outside the scope of existing theoretical frameworks, hindering the development of Lipschitz-continuous models in acoustic applications. In particular, despite their widespread adoption, DNNs that separately process the magnitude and phase of complex-valued signals cannot be Lipschitz continuous under existing frameworks. In this paper, to address this limitation, we establish a theoretical foundation for constructing amplitude modifiers (AMs), a class of DNN architectures that operate solely on the magnitude of a complex-valued input, with provable Lipschitz continuity. Specifically, we derive a necessary and sufficient condition for an AM to be Lipschitz continuous and propose LipsAMs (Lipschitz-continuous AMs) corresponding to common architectures for audio signals, including time-frequency masking. Furthermore, we develop an efficient framework for evaluating their Lipschitz constants and analytically derive these constants for some of the proposed architectures. As an application, we propose CoReM-LipsAM (Controlled Residual Maps via LipsAM) for plug-and-play (PnP) audio signal recovery, integrating a DNN as a data-driven prior within a model-based signal processing algorithm. The convergence of the obtained PnP algorithm is structurally guaranteed by the CoReM-LipsAM architecture and empirically validated through speech dereverberation experiments.

发表机构

  • Tokyo University of Agriculture and Technology(东京农工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑