arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NeuMark-Native:通过充分利用神经音频编解码器潜在空间实现鲁棒的文本到语音原生水印

NeuMark-Native: Robust Text-to-Speech-Native Watermarking Through Full Utilization of Neural Audio Codec Latent Space

Annan Wu, Wen-Chin Huang, Tomoki Toda

arXiv 2610.05215首次发表:更新:

发表机构

Nagoya University(名古屋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

NeuMark-Native提出一种TTS原生水印框架,在神经编解码器潜在空间中嵌入载荷,提升水印在DSP和编解码器攻击下的持久性,同时保持语音质量。

AI 中文摘要

语音水印为合成语音提供了主动的可追溯性,然而现有的大多数模型仅在文本到语音(TTS)合成之后,通过向生成的波形添加水印扰动来运行。这种事后设计使水印成为可省略或可绕过的外部步骤,并将水印限制在浅层波形表示中。我们提出了NeuMark-Native,一种面向基于神经编解码器合成的TTS原生水印框架。它在波形解码之前,将载荷信息嵌入到每个生成的编解码器潜在层中,从而提高了水印在下游数字信号处理(DSP)和神经编解码器再合成下的持久性。NeuMark-Native保持预训练的TTS模型和神经编解码器冻结,仅优化生成编解码器令牌上的水印模块。在两个语料库上,针对11种DSP攻击和9种神经编解码器攻击的实验表明,该框架在保持自然度、可懂度和接近合成语音的语音质量的同时,实现了鲁棒的水印检测。

英文摘要

Speech watermarking offers proactive traceability for synthetic speech, yet most existing models operate only after text-to-speech (TTS) synthesis by adding a watermark perturbation to the generated waveform. This post-hoc design leaves watermarking as an external step that can be omitted or bypassed and restricts the watermark to a shallow waveform representation. We propose NeuMark-Native, a TTS-native watermarking framework for neural codec-based synthesis. It embeds payload information into every generated codec-latent layer before waveform decoding, improving watermark persistence under downstream digital signal processing (DSP) and neural codec resynthesis. NeuMark-Native keeps the pretrained TTS model and the neural codec frozen, while optimizing only the watermark modules on generated codec tokens. Experiments on two corpora under 11 DSP attacks and 9 neural-codec attacks demonstrate robust watermark detection while preserving naturalness, intelligibility, and speech quality close to synthetic speech.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑