arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经音频水印在语音增强下的脆弱性

The Vulnerability of Neural Audio Watermarks under Speech Enhancement

Xincong Zhong, Shengyao Wang, Lingfeng Yao, Yihang Bao, Jinze Yu, Miao Pan, Jiang Liu

arXiv 2609.29040首次发表:更新:

发表机构

Waseda University; University of Houston(早稻田大学; 休斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将高斯噪声与语音增强模型级联作为黑盒攻击,发现生成式语音增强能有效去除六种神经音频水印,揭示其脆弱性并呼吁鲁棒性评估。

AI 中文摘要

神经音频水印越来越多地部署在商业语音生成系统中,以使AI生成的语音可追溯,然而其鲁棒性主要是在传统信号失真条件下研究的。由于水印可以被视为添加到语音信号中的不可感知噪声,一个自然的问题是,语音增强(SE)作为一种去噪模型,能否去除水印。在本文中,我们将高斯噪声与SE模型级联,作为一种黑盒水印去除攻击,覆盖判别式和生成式两种SE范式,针对六种神经水印:AudioSeal、WavMark、SilentCipher、Timbre、Perth和AlignMark。实验结果表明,所提出的攻击在水印去除方面显著优于现有的神经重合成方法。特别是,我们发现生成式SE在去噪的同时重建语音的谐波区域,对水印具有高度破坏性。这些发现表明,SE对当前的音频水印方法构成了严重威胁,我们呼吁在水印设计中纳入SE感知的鲁棒性评估。

英文摘要

Neural audio watermarks are increasingly deployed in commercial speech generation systems to make AI-generated speech traceable, yet their robustness has been studied mainly under conventional signal distortions. Since a watermark can be regarded as imperceptible noise added to the speech signal, a natural question is whether speech enhancement (SE), as a denoising model, can remove it. In this paper, we cascade Gaussian noise with SE models as a black-box watermark removal attack, covering both discriminative and generative SE paradigms, against six neural watermarks: AudioSeal, WavMark, SilentCipher, Timbre, Perth, and AlignMark. Experimental results show that the proposed attack significantly outperforms existing neural re-synthesis methods in watermark removal. In particular, we find that generative SE, which reconstructs the harmonic regions of speech while denoising, is highly destructive to watermarks. These findings show that SE poses a serious threat to current audio watermarking methods, and we call for SE-aware robustness evaluation in watermark design.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑