arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AudioNoisePrints:基于流匹配TTS空间相关性的无模型音频水印

AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS

Timothy Tin-Long, Jian Zhu, Aidan Pine, Mengzhe Geng

arXiv 2608.22186首次发表:更新:

发表机构

National Research Council Canada; University of British Columbia(加拿大国家研究委员会; 不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出无训练的AudioNoisePrints音频水印流水线,利用流匹配TTS的噪声-输出相关性实现水印,在强增强下优于AudioSeal,可适配多种TTS模型。

AI 中文摘要

我们提出了AudioNoisePrints,一种用于流匹配和扩散TTS模型的无训练水印流水线,该流水线在推理过程中仅需极少的额外计算,且无需重新训练TTS模型或降低生成质量。我们利用了扩散和流匹配模型中初始高斯噪声与生成输出之间存在强相关性这一特性,因此可通过初始噪声与生成输出之间的简单余弦相关性来实现水印。此外,我们在其上方训练了一个轻量级检测器以应对更强的增强操作。我们的方法在强增强条件下的音频水印基准AudioSeal上表现更优。我们在F5TTS及其他TTS和声码器模型上进行了实验,得出这些模型均表现出相似的空间相关性特性,这表明我们的水印方案未来可应用于更多流匹配TTS模型甚至声码器。

英文摘要

We present AudioNoisePrints, a training-free watermarking pipeline for flow matching and diffusion TTS models, which requires minimal extra computation during inference and does not require retraining the TTS model or reducing the generation quality. We exploited the fact that there are strong correlations between the initial Gaussian noises and the generated outputs in diffusion and flow matching models, such that a simple cosine correlation between the initial noise and the generated output can be used to perform watermaking. Moreover, we train a lightweight detector on top for more aggressive augmentations. Our method outperforms AudioSeal, a strong baseline for audio watermarking under strong augmentations. We experimented on F5TTS and other TTS and vocoder models, and concluded that they all exhibit similar spatial correlation properties, suggesting our watermarking scheme can be used for more flow-matching TTS models and even vocoders in the future.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑