发表机构
Imperial College London; Queen Mary University of London; Johns Hopkins University; Technical University of Munich(帝国理工学院; 伦敦玛丽女王大学; 约翰斯·霍普金斯大学; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出PIPDP协议区分语音深度伪造的可感知与不可感知被动指纹,实验显示不可感知指纹的归因线索更持久可靠,为语音深度伪造归因提供了新的诊断方法。
AI 中文摘要
被动指纹(生成器自然留下的内在痕迹)已被证明可用于语音深度伪造检测中的归因,但它们的持久性、可复现性和内容独立性仍未得到验证。此外,现有研究未区分可感知与不可感知指纹,尽管二者对归因可靠性的影响差异极大。可感知指纹(如情感表达)由感知质量目标塑造,可能随模型更新而变化;不可感知指纹未被当前训练目标明确优化,且在现有数据集设计或训练策略中很少被考虑,因为它们对下游应用的影响有限。因此,我们提出可感知-不可感知被动指纹诊断协议(PIPDP),以定义并分别分析这两种指纹类型。PIPDP包含三项互补分析:通过残差能量、可复现性和显著性分析进行的多证据指纹验证,保留音频质量的感知透明扰动,以及无需模型重训练即可修改可感知指纹的提示驱动情感变化。针对10种语音生成器和3种归因检测器的实验表明,不可感知指纹提供了持久的归因线索。感知透明扰动在HiggsAudioV3上使归因准确率降低多达48.2%,而情感驱动的变化对归因几乎无影响,在使用w2v-bert-MLP的CosyVoice2上,不同情感间的准确率变化仅约1.0%。这些结果表明,不可感知指纹更适合用于可靠归因。
英文摘要
Passive fingerprints (intrinsic traces naturally left by generators) have been shown to enable attribution in speech deepfake detection, yet their persistence, reproducibility, and content-independence remain unverified. Moreover, no prior work distinguishes perceptible from imperceptible fingerprints, although the two have very different implications for attribution reliability. Perceptible fingerprints, such as emotional expression, are shaped by perceptual quality objectives and may change across model updates, whereas imperceptible fingerprints are not explicitly optimised by current training objectives and are rarely considered in existing dataset design or training strategies, as they have limited influence on downstream applications. We therefore propose a Perceptible-Imperceptible Passive-fingerprint Diagnostic Protocol (PIPDP) to define and separately analyze these two fingerprint types. PIPDP comprises three complementary analyses: multi-evidence fingerprint verification through residual-energy, reproducibility, and saliency analyses, perceptually transparent perturbations preserving audio quality, and prompt-driven emotion change that modifies perceptible fingerprints without model retraining. Experiments across ten speech generators and three attribution detectors show that imperceptible fingerprints provide persistent attribution cues. Perceptually transparent perturbations reduce attribution accuracy by up to 48.2\% on HiggsAudioV3, whereas emotion-driven changes leave attribution largely unchanged, with only about a 1.0\% accuracy variation across emotions on CosyVoice2 using w2v-bert-MLP. These results suggest that imperceptible fingerprints are more reliable for trustworthy attribution.