发表机构
East China Normal University(华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自回归音频生成水印易受编解码器攻击的问题,提出MARC多比特水印方法,结合固有令牌表示与编解码器混淆模式,在语音、对话、音乐生成任务中实现高比特提取准确率与攻击鲁棒性。
AI 中文摘要
生成音频现已应用于诸多场景,因此在其分发及信号处理后需验证来源。对于自回归音频生成,该任务极具挑战性,因为编解码器处理会改变从波形恢复的令牌序列,这类变化会降低水印检测和载荷解码的可靠性。现有方法要么利用固有令牌表示,要么利用变换导致的替换模式来构建令牌组,导致固有令牌关系与编解码器诱导的替换被分开建模;此外,多数方法仅支持零比特检测,只能判断水印是否存在,无法区分不同的生成输出。本文提出MARC,一种用于自回归音频生成的多比特生成式水印方法。MARC将固有令牌表示与通过重令牌化及多种编解码器获得的混淆模式相结合,形成感知编解码器的令牌聚类空间;在该空间内,采用载荷驱动的聚类调度来嵌入多比特水印,检测和载荷解码则在重令牌化后的观测结果上执行。针对语音、对话及音乐生成的实验表明,MARC在未修改的含水印音频上实现了平均97.3%的比特提取准确率,且在多种编解码器攻击下仍可提取水印,同时对覆盖攻击也表现出鲁棒性。
英文摘要
Generated audio is now used in a range of applications, creating a need to verify its origin after distribution and signal processing. This task is particularly challenging for autoregressive audio generation because codec processing can alter the token sequence recovered from the waveform. Such changes reduce the reliability of watermark detection and payload decoding. Existing methods construct token groups using either intrinsic token representations or substitution patterns caused by transformations. As a result, intrinsic token relationships and codec induced substitutions are modeled separately. In addition, most methods support only zero bit detection. They can determine whether a watermark is present but cannot distinguish individual generated outputs. We propose \textbf{MARC}, a multi-bit generative watermarking method for autoregressive audio generation. MARC integrates intrinsic token representations with confusion patterns obtained through retokenization and multiple codecs, forming a codec-aware token-cluster space. Within this space, payload-driven cluster scheduling is used to embed a multi-bit watermark, while detection and payload decoding are performed on retokenized observations. Experiments on speech, dialogue, and music generation show that MARC achieves an average of 97.3\% bit extraction accuracy on unmodified watermarked audio and the watermark can still be extracted under diverse codec attacks. MARC also demonstrates robustness to overwriting attacks.