CRAW:编解码器鲁棒音频水印
CRAW: Codec Robust Audio Watermarking
浏览论文内容
中文总结 AI 辅助
针对现有音频水印方法在实际变换下鲁棒性不足的问题,提出CRAW框架,结合多种技术提升对神经编解码器等的鲁棒性,同时保持感知质量,达最优性能。
中文摘要 AI 辅助
近期生成式语音模型的进展使得区分真实音频与合成音频愈发困难,催生了新的欺诈与虚假信息形式。音频水印通过在生成语音中嵌入不可感知的信号,后续可检测该信号以验证来源,是一种颇具前景的防御手段。然而,现有事后水印方法在神经编解码器和去噪器下失效,而这些变换是实际存储、传输和处理过程中常规应用的操作,严重限制了其实际应用。本文提出CRAW,一种编解码器鲁棒音频水印框架,可在保持高感知质量的同时,共同提升对神经重合成的鲁棒性。CRAW结合了感知失真感知训练、基于注意力的池化机制、推理时感知掩码以及纠错码,以恢复鲁棒训练过程中损失的保真度。实验表明,CRAW在保持与现有事后水印方法相当的感知质量的同时,实现了对神经编解码器、去噪器和声码器的最先进鲁棒性。代码可在指定URL获取。
英文摘要
Recent advances in generative speech models have made it increasingly difficult to distinguish authentic from synthetic audio, enabling new forms of fraud and misinformation. Audio watermarking offers a promising defense by embedding an imperceptible signal into generated speech that can later be detected to verify its provenance. However, recent studies have shown that existing post-hoc watermarking methods fail under neural codecs and denoisers, transformations routinely applied during real-world storage, transmission, and processing, severely limiting their practical utility. Here we introduce CRAW, a codec-robust audio watermarking framework that jointly improves robustness against neural re-synthesis while maintaining high perceptual quality. CRAW combines distortion-aware training with an attention-based pooling mechanism, inference-time perceptual mask- ing, and an error-correcting code to recover the fidelity lost during robust training. Experiments demonstrate that CRAW achieves state-of-the-art robustness against neural codecs, denoisers, and vocoders while maintaining perceptual quality comparable to existing post-hoc watermarking methods. The code is available at https://github.com/DavidC1212/craw.
发表机构
- Bar-Ilan University(巴伊兰大学)
- NVIDIA(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。