arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00662eess.AS

Silence-the-Mimic:加速针对语音克隆的不可感知扰动生成

Silence-the-Mimic: Accelerating Imperceptible Perturbation Generation Against Voice Cloning

  • The University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

Runqiu Xu

AI总结:

针对语音克隆的隐私风险,提出频域心理声学约束的快速对抗扰动生成方法STM,在严格保证不可感知性的同时,实现高达45.3倍加速和更优保护性能。

AI中文摘要:

基于深度神经网络的语音转换(VC)和文本到语音(TTS)模型发展迅速,能够以最少输入数据实现逼真的语音克隆。这种能力引发了对未经授权克隆说话者身份及其相关隐私和安全风险的严重担忧。现有的不可感知对抗性保护方法依赖于质量控制损失,这些损失对超参数调整高度敏感,且由于优化过程冗长而计算成本高昂。为解决这些局限性,我们提出了一种快速保护方法,在心理声学掩蔽约束下,于频域中生成感知受限的扰动。我们的方法在对抗训练期间严格强制执行可感知性界限,消除了迭代质量平衡的需要,并显著降低了计算成本。在多个最先进的VC和TTS模型上的实验表明,STM实现了具有竞争力或更优的保护性能,同时具有显著更好的感知质量,并且相比现有白盒基线实现了高达45.3倍的加速。这些结果证明了具有感知约束的频域扰动作为保护免受语音克隆侵害的实用范式的有效性。

英文摘要:

Deep neural network-based Voice Conversion (VC) and Text-to-Speech (TTS) models have rapidly advanced, enabling realistic voice cloning with minimal input data. Such capabilities raise serious concerns over unauthorized cloning of speaker identities and the associated privacy and security risks. Current imperceptible adversarial protection methods rely on quality control losses that are highly sensitive to hyperparameter tuning and computationally expensive due to lengthy optimization. To address these limitations, we propose a fast protection method that generates perceptually constrained perturbations in the frequency domain under a psychoacoustic masking-based constraint. Our approach strictly enforces perceptibility bounds during adversarial training, eliminating the need for iterative quality balancing and significantly reducing computational cost. Experiments on multiple state-of-the-art VC and TTS models show that STM achieves competitive or superior protection performance with substantially better perceptual quality and up to $45.3\times$ speedup over existing white-box baselines. These results demonstrate the effectiveness of frequency-domain perturbations with perceptual constraints as a practical paradigm for protecting against voice cloning.

补充信息

↑