发表机构
Korea University; Inha University(韩国大学; 仁荷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人工智能音乐生成中水印溯源问题,现有方法多为事后处理且易受攻击。本文提出MusicMark,在生成时将水印嵌入语义潜在空间,通过水印适配器和联合目标训练,实验证明其在多种攻击下优于基线,翻唱攻击中也更具鲁棒性。
AI 中文摘要
随着商业平台的发展,人工智能音乐生成迅速进步,这就需要可靠的水印来进行溯源和归属认定。然而,现有的音频水印研究主要集中在语音上,由于音乐结构复杂、声学纹理丰富,将面向语音的方法应用于音乐具有挑战性。大多数现有方法是事后处理的,在生成后添加难以察觉的扰动,而不是将水印作为内容的一部分嵌入。这使得它们在变换下很脆弱,尤其容易受到神经编解码器重新合成的影响,因为这种重新合成可能会丢弃难以察觉的残留信号。此外,由于生成和水印是解耦的,水印步骤可能会被绕过或省略,削弱了溯源保证。为了解决这些问题,我们提出了MusicMark,据我们所知,它是第一个用于音乐的生成式水印框架。具体来说,MusicMark在生成过程中将水印消息嵌入语义潜在空间,将水印作为音乐内容的一部分,并确保对各种攻击具有鲁棒性,特别是对神经编解码器重新合成。为此,我们在基于扩散的生成模型中引入了一个水印适配器,以在去噪步骤中嵌入水印消息。适配器和检测器通过联合目标进行训练,通过约束带水印的潜在向量接近其无水印的参考潜在向量来保持保真度,同时通过攻击增强来提高鲁棒性。实验表明,MusicMark在包括神经编解码器重新合成在内的各种攻击中显著优于事后处理基线,同时保持了可比的生成质量。我们还引入了一种翻唱攻击,在保留音乐内容的同时转换歌声,并表明MusicMark比事后处理方法更具鲁棒性。
英文摘要
AI music generation has rapidly advanced alongside commercial platforms, raising the need for reliable watermarking for provenance and attribution. However, existing audio watermarking research has largely focused on speech, and applying speech-oriented methods to music is challenging due to music's complex structure and rich acoustic texture. Most existing methods are post-hoc, adding imperceptible perturbations after generation rather than embedding watermarks as part of the content. This makes them fragile under transformations and especially vulnerable to neural codec re-synthesis, which can discard imperceptible residual signals. Moreover, since generation and watermarking are decoupled, the watermarking step can be bypassed or omitted, weakening provenance guarantees. To address these issues, we propose MusicMark, which, to the best of our knowledge, is the first generative watermarking framework for music. Specifically, MusicMark embeds watermark messages into the semantic latent space during generation, incorporating the watermark as part of the musical content and ensuring robustness against diverse attacks, particularly neural codec re-synthesis. To this end, we introduce a watermark adapter into a diffusion-based generation model to embed watermark messages across denoising steps. The adapter and detector are trained with a joint objective that preserves fidelity by constraining watermarked latents close to their unwatermarked reference latents, while improving robustness through attack augmentations. Experiments demonstrate that MusicMark substantially outperforms post-hoc baselines across diverse attacks including neural codec re-synthesis, while maintaining comparable generation quality. We further introduce a cover-song attack, converting the singing voice while preserving musical content, and show that MusicMark remains more robust than post-hoc methods.
CommentsSubmitted to IEEE Transactions on Information Forensics and Security