发表机构
National Institute of Informatics(情报信息研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对部分深度伪造语音检测难题,本文提出一种无需训练的主动防御方法,通过将自嵌入策略与现有音频隐写术结合,借助编解码器修复实现检测,可与被动防御互补且数据效率高。
AI 中文摘要
仅合成或操纵话语有限片段的部分深度伪造语音,对现有深度伪造检测系统构成重大挑战。随着伪造区域比例降低,被动检测器可靠性急剧下降,准确检测与修复仍具挑战性。本文从新视角重新审视音频隐写术,提出将其用作针对部分深度伪造音频的主动防御手段。具体而言,我们采用自嵌入策略:干净语音信号嵌入自身的压缩表示,支持事后提取参考内容。我们展示了如何重新利用现有音频隐写术方法,通过基于编解码器的修复来支持部分深度伪造检测。在基准数据集上的实验表明,所提方法可与被动防御形成互补。值得注意的是,该方法无需任何训练,为部分深度伪造检测提供了一种鲁棒且数据高效的替代方案。
英文摘要
Partial deepfake speech, where only limited segments of an utterance are synthesized or manipulated, poses a significant challenge to existing deepfake detection systems. As the proportion of spoofed regions decreases, passive detectors become increasingly unreliable, and accurate detection and restoration remain challenging. In this paper, we revisit audio steganography from a new perspective and propose its use as a proactive defense against partially deepfaked audio. In particular, we consider a self-embedding strategy in which a clean speech signal embeds a compressed representation of itself, enabling post-hoc extraction of reference content. We demonstrate how existing audio steganography methods can be repurposed to support detection of partial deepfakes through codec-based restoration. Experiments on a benchmark dataset show that the proposed approach complements passive defenses. Remarkably, the proposed method operates without any training, providing a robust and data-efficient alternative for partial deepfake detection.
Comments6 pages; 4 figures; 1 tables; accepted at Interspeech 2026; audio samples available at https://nii-yamagishilab.github.io/self-embedding-audio-stego-demo-pages/