发表机构
CITI, Academia Sinica; IIS, Academia Sinica(资讯科技创新中心,中央研究院; 资讯科学研究所,中央研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对合成语音的相关问题,提出免训练的SSTMark语音水印框架,通过文本水印在语义级别操作,将水印信息编码到语音语义内容中并检测。实验表明其平均鲁棒性最强,相比基线在不同编辑下检测率有显著提升。
AI 中文摘要
随着语音生成模型越来越逼真且广泛可用,对合成语音的滥用、归属和治理的担忧不断增加。水印提供了一种使合成语音可追溯和可验证的实用方法。大多数现有语音水印方法将水印信息嵌入到信号级表示中,如波形或频谱图。在足够强的失真下,嵌入的水印可能会被削弱或破坏,导致检测能力下降。本文提出了SSTMark,一种通过文本水印在语义级别运行的免训练语音水印框架。与传统的信号级水印方法不同,SSTMark将水印信息编码到生成语音所传达的语义内容中,并从恢复的语言内容中检测水印。在AudioMarkBench上的实验表明,SSTMark具有最强的平均鲁棒性。在固定误报率为1%的情况下,与最先进的基线相比,SSTMark在信号处理编辑和压缩编辑上的平均检测率分别提高了4.6%和16.9%。
英文摘要
As speech generation models become increasingly realistic and widely accessible, concerns about the misuse, attribution, and governance of synthetic speech continue to grow. Watermarking provides a practical way to make synthesized speech traceable and verifiable. Most existing speech watermarking methods embed watermark information into signal-level representations, such as waveforms or spectrograms. Under sufficiently strong distortions, the embedded watermark may be weakened or destroyed, leading to degraded detectability. In this paper, we propose SSTMark, a training-free speech watermarking framework that operates at the semantic level through text watermarking. Unlike conventional signal-level watermarking methods, SSTMark encodes watermark information into the semantic content conveyed by generated speech, and detects the watermark from the recovered linguistic content. Experiments on AudioMarkBench demonstrate that SSTMark exhibits the strongest average robustness. Compared with the state-of-the-art baselines at a fixed false positive rate of 1\%, SSTMark improves the average detection rate by 4.6\% and 16.9\% on signal-processing edits and compression edits, respectively.