KeyBound:用于语音溯源的有密钥且宿主绑定的学习型音频水印
KeyBound: Keyed and Host-Bound Learned Audio Watermarking for Speech Provenance
浏览论文内容
中文总结 AI 辅助
KeyBound提出有密钥且宿主绑定的学习型音频水印,通过密钥掩蔽和宿主调制载体实现安全溯源,在LibriSpeech上优于现有基线,并抵抗重合成和移植。
中文摘要 AI 辅助
音频水印是一种将合成语音归因于其来源的主动方法。学习型音频水印通常通过在固定的信号失真目录(如噪声、压缩、滤波和重采样)下恢复载荷来评判。该测试对于溯源是必要的,但不够充分。作为来源证据的标记不应被未授权方读取,不应可转移到无关音频,也不应在现代生成模型重新合成录音时消失。我们提出了KeyBound,一种学习型音频水印,它恢复了经典水印提供而学习方案搁置的两个要素:秘密密钥和宿主感知的载体。KeyBound用秘密密钥掩蔽载荷,并通过由宿主的冻结频谱表示调制的载体嵌入掩蔽比特,因此密钥控制载荷访问,而宿主条件载体抵抗直接移植。密钥无关的存在性检测头允许任何方检测标记,而只有密钥持有者才能读取其归属,并且在单样本均匀性假设下,错误密钥解码通过我们验证规则的概率至多为$2.1\ imes10^{-3}$。在LibriSpeech上,与WavMark、AudioSeal和Timbre相比,KeyBound在使所有基线检测失效的频谱去噪器下保持1.00的检测准确率和0.98的比特准确率,无密钥时解码为随机水平,并拒绝移植的载体。检测进一步转移到保留的DAC和BigVGAN重合成,尽管精确载荷恢复有所下降。因此,语音溯源更适合被表述为一个有密钥的、宿主绑定的归因问题,而不是在预先固定的信号失真目录下恢复载荷。
英文摘要
Audio watermarking is a proactive route to attributing synthetic speech to its source. Learned audio watermarks are typically judged by payload recovery after a fixed catalog of signal distortions such as noise, compression, filtering, and resampling. That test is necessary but not sufficient for provenance. A mark offered as evidence of origin should not be readable by an unauthorized party, should not be transferable to unrelated audio, and should not vanish when the recording is re-synthesized by a modern generative model. We present KeyBound, a learned audio watermark that restores the two ingredients classical watermarking supplied and learned schemes set aside, a secret key and a host-aware carrier. KeyBound masks the payload with a secret key and embeds the masked bits through a carrier modulated by a frozen spectral representation of the host, so the key governs payload access while the host-conditioned carrier resists direct transplantation. A key-independent presence head lets any party detect a mark, whereas only a key holder reads its attribution, and under the single-sample uniformity assumption a wrong-key decode clears our verification rule with probability at most $2.1\times10^{-3}$. On LibriSpeech against WavMark, AudioSeal, and Timbre, KeyBound holds 1.00 detection accuracy and 0.98 bit accuracy under a spectral denoiser that costs every baseline its detection, decodes at chance without the key, and rejects transplanted carriers. Detection further transfers to held-out DAC and BigVGAN re-synthesis, though exact payload recovery degrades. Speech provenance is thus better posed as a keyed, host-bound attribution problem than as the recovery of a payload under a catalog of signal distortions fixed in advance.
发表机构
- University of New South Wales(新南威尔士大学)
- National University of Singapore(新加坡国立大学)
- Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。