arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NeuMark:编解码器潜空间中的神经编解码器重合成鲁棒音频水印

NeuMark: Neural Codec Resynthesis-Robust Audio Watermarking in the Codec Latent Space

Annan Wu, Wen-Chin Huang, Tomoki Toda

arXiv 2609.25719首次发表:更新:

发表机构

Nagoya University(名古屋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

NeuMark在编解码器潜空间嵌入水印,利用交叉注意力在RVQ层注入16位消息,显著提升神经编解码器重合成下的鲁棒性,兼顾检测与恢复。

AI 中文摘要

音频水印对于追踪生成的语音日益重要。已有多种音频水印方法被提出,将水印嵌入到波形、音色特征或潜在表示等不同域中,以使嵌入的水印对传统数字信号处理(DSP)攻击具有鲁棒性。另一方面,现代神经编解码器引入了不同于DSP攻击的威胁:它们通过量化的声学表示重合成语音,并可能移除与编解码器保留结构不对齐的嵌入水印证据。在本文中,我们提出NeuMark,一种编解码器潜空间音频水印框架,将水印证据嵌入到SpeechTokenizer声学标记中,以应对这种重合成威胁。NeuMark使用交叉注意力在残差向量量化(RVQ)层之间注入16位消息,将水印分布到与编解码器对齐的潜结构上。实验结果表明,NeuMark在神经编解码器重合成下显著提高了鲁棒性,同时支持水印检测和消息恢复。我们还分析了重建参考透明性与原始参考鲁棒性之间的权衡。

英文摘要

Audio watermarking is increasingly important for tracing generated speech. Several audio watermarking methods have been proposed to embed the watermark in various domains, such as waveform, timbre feature, or latent representations, for making the embedded watermark robust against traditional digital signal processing (DSP) attacks. On the other hand, modern neural codecs introduce a different threat from DSP attacks: they resynthesize speech through quantized acoustic representations and can remove the embedded watermark evi- dence that is not aligned with codec-preserved structure. In this paper, we propose NeuMark, a codec-latent audio watermarking framework that embeds watermark evidence into SpeechTok- enizer acoustic tokens to address this resynthesis threat. NeuMark uses cross-attention to inject a 16-bit message across residual vector quantization (RVQ) layers, distributing the watermark over codec-aligned latent structure. Experimental results show that NeuMark substantially improves robustness under neural- codec resynthesis while supporting both watermark detection and message recovery. We also analyze the trade-off between reconstruction-referenced transparency and original-referenced robustness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑