arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

忘记你所忘记的:零样本语音合成中的说话人遗忘以防止重新识别

Forget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech

Hyoeun Kim, Yujun Lee, Kyuhong Shim

arXiv 2609.27399首次发表:更新:

发表机构

Sungkyunkwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对零样本语音合成中的说话人身份遗忘问题,提出GUARD框架,结合门控与激活引导,在CosyVoice2上将重新识别准确率从73.5%降至0.5%,并强调重新识别作为补充评估标准。

AI 中文摘要

最近的零样本文本到语音(ZS-TTS)系统仅需几秒钟的参考语音即可高保真地重现说话人的声音,这引发了对未经授权的声音克隆和冒充的担忧。说话人身份遗忘最近作为一种方法出现,旨在选择性地抑制对选择退出的说话人的这种能力,同时保留对其他说话人的合成能力。尽管现有方法降低了说话人相似度,但防止重新识别往往面临语音质量的严重下降。受此观察启发,我们提出了GUARD,一个轻量级的说话人身份遗忘框架,该框架将学习到的说话人门控与在冻结的TTS骨干上的说话人无关激活引导相结合。引导向量通过组相对奖励优化进行优化,以将输出从被遗忘说话人转向群体水平的冒名顶替者相似度,同时保留可懂度和语音自然度。在CosyVoice2上,GUARD将被遗忘说话人的相似度从0.541降至0.103,在150个说话人的图库中的重新识别准确率从73.5%降至0.5%,同时保留了对保留说话人的重现。结果表明,仅降低相似度可能无法完全表征成功的说话人身份遗忘,并强调重新识别作为其评估的补充标准。

英文摘要

Recent zero-shot text-to-speech (ZS-TTS) systems can reproduce a speaker's voice with high fidelity from only a few seconds of reference speech, raising concerns over unauthorized voice cloning and impersonation. Speaker identity unlearning has recently emerged as an approach to selectively suppress this capability for speakers who opt out while preserving synthesis capability for other speakers. Although existing approaches reduce speaker similarity, preventing re-identification often faces severe degradation of speech quality. Motivated by this observation, we propose GUARD, a lightweight speaker identity unlearning framework that combines a learned speaker gate with speaker-agnostic activation steering on a frozen TTS backbone. The steering vectors are optimized using group-relative reward optimization to shift outputs from forget speakers toward population-level impostor similarity while preserving intelligibility and speech naturalness. On CosyVoice2, GUARD reduces forget-speaker similarity from 0.541 to 0.103 and re-identification accuracy in a 150-speaker gallery from 73.5% to 0.5%, while preserving retain-speaker reproduction. The results demonstrate that similarity reduction alone may not fully characterize successful speaker identity unlearning and highlight re-identification as a complementary criterion for its evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑