arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TTS-Guard:通过自适应对抗说话人配对指纹实现文本到语音模型的黑盒所有权验证

TTS-Guard: Black-Box Ownership Verification of Text-to-Speech Models via Adaptive Adversarial Speaker-Pair Fingerprints

Xubin Yue, Zhenhua Xu, Zhebo Wang, Mengting Li, Zijie Zhou, Wenpeng Xing, Dezhang Kong, Meng Han

arXiv 2609.23729首次发表:更新:

发表机构

Zhejiang University; Binjiang Institute of Zhejiang University; Hangzhou International Innovation Institute, Beihang University; China University of Petroleum (Beijing)(浙江大学; 浙江大学滨江研究院; 北京航空航天大学杭州国际创新研究院; 中国石油大学(北京))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TTS-Guard提出基于对抗性说话人配对指纹的黑盒所有权验证框架,通过双嵌入空间和自适应课程优化,在五个TTS系统上实现96.4%指纹成功率,且对多种攻击和修改保持鲁棒性。

AI 中文摘要

零样本文本到语音(TTS)模型的快速成熟已将高质量语音克隆转变为广泛可用的能力,引发了对专有语音模型未经授权复制、微调和转售的严重担忧。然而,TTS的所有权验证在很大程度上仍然是一个开放问题:语音是连续波形,其扰动很容易被常规信号处理破坏,并且人类听觉系统施加的感知预算比视觉严格得多。我们提出了TTS-Guard,一种基于对抗性说话人配对指纹的TTS模型黑盒所有权验证框架。TTS-Guard(i)在双嵌入空间中选择关键说话人配对,以实现架构无关的隐蔽性;(ii)通过覆盖微调、剪枝、量化和蒸馏的影子模型自适应课程优化扰动;(iii)将黑盒查询聚合为校准的验证置信度分数。在五个主流TTS系统上,TTS-Guard在假阳性率为5.8%的情况下达到了平均96.4%的指纹成功率,同时保持了可懂性和自然度。该指纹对十种音频攻击、六种模型修改和两种最先进的对抗性净化器仍然有效。

英文摘要

The rapid maturation of zero-shot Text-to-Speech (TTS) models has turned high-quality voice cloning into a widely available capability, raising acute concerns over unauthorised replication, fine-tuning and resale of proprietary speech models. Yet ownership verification for TTS remains largely open: speech is a continuous waveform whose perturbations are easily destroyed by routine signal processing, and the human auditory system imposes a much tighter perceptual budget than vision. We present \textbf{TTS-Guard}, a black-box ownership verification framework for TTS models built on \emph{adversarial speaker-pair fingerprints}. TTS-Guard(i) selects key speaker pairs in a \emph{dual} embedding space for architecture-agnostic stealth;(ii) optimises a perturbation through an \emph{adaptive curriculum} of shadow models covering fine-tuning, pruning, quantisation and distillation; and (iii) aggregates black-box queries into a calibrated \emph{Verification Confidence Score}. On five mainstream TTS systems, TTS-Guard reaches an average Fingerprint Success Rate of $96.4\%$ at a False Positive Rate of $5.8\%$, while preserving intelligibility and naturalness. The fingerprint remains effective against ten audio attacks, six model modifications, and two state-of-the-art adversarial purifiers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑