发表机构
East China Normal University; Huaqiao University(华东师范大学; 华侨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对TTS合成语音在重建攻击下溯源难的问题,提出Thrive多比特生成式水印框架,通过同步注入与逐位可靠性选择,在保持保真度下实现87.6%恢复准确率及万级身份归属。
AI 中文摘要
现代文本转语音(TTS)系统日益大规模地为不同用户生成合成语音。这种场景要求内容级别的溯源能力,以验证发布语音的来源并将其归属于请求用户,而生成式水印可以通过将多位标识符直接嵌入合成语音来支持这一需求。然而,语音一旦发布,在分发和编辑过程中可能经历异构的学习变换,这些变换的重建目标可以在保留语音实用性的同时,对水印的可恢复性产生不同影响。我们发现,没有任何单一的水印载体能在所有重建模型下保持一致的可靠性,因为其存活性共同取决于嵌入结构、重建机制和观测表示。为此,我们提出了Thrive,一个面向现代自回归TTS的多位生成式语音水印框架,覆盖了在重建攻击下的离散令牌和连续表示生成。具体而言,Rise将水印注入中间表示,并与其在后续生成中的持续集成同步,而Care则通过逐位可靠性选择结合波形和频谱专家。在两种自回归范式上的实验表明,Thrive保持了合成保真度,在重建攻击下实现了87.6%的平均恢复准确率,并支持在多达10,000个身份的候选集中进行来源归属。
英文摘要
Modern TTS systems increasingly generate synthetic speech at scale for diverse users. This setting calls for content-level provenance that can verify the origin of released speech and attribute it to the requesting user, which generative watermarking can support by embedding multi-bit identifiers directly into synthesized speech. Once released, however, speech may undergo heterogeneous learned transformations during distribution and editing, with reconstruction objectives that can preserve speech utility while affecting watermark recoverability differently. We find that no single watermark carrier remains consistently reliable across reconstruction models, as its survival depends jointly on the embedded structure, reconstruction mechanism, and observation representation. To this end, we propose Thrive, a multi-bit generative speech watermarking framework for modern autoregressive TTS, covering both discrete-token and continuous-representation generation under reconstruction attacks. Specifically, Rise synchronizes watermark injection into intermediate representations with its continued integration into subsequent generation, while Care combines waveform and spectral experts using bit-wise reliability selection. Experiments on both autoregressive paradigms show that Thrive preserves synthesis fidelity, achieves 87.6% average recovery accuracy under reconstruction attacks, and supports source attribution over candidate sets of up to 10,000 identities.