语言正交化用于零样本跨语言音频深度伪造检测
Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection
浏览论文内容
中文总结 AI 辅助
针对音频深度伪造检测的跨语言泛化问题,提出语言正交化方法,通过移除自监督语音模型中的语言相关变异,在多种语言和骨干网络上一致降低等错误率。
中文摘要 AI 辅助
音频深度伪造检测器需要迁移到训练中未出现的语言,因为多语言语音合成的发展速度超过了标注的反欺骗资源。尽管检测器越来越依赖自监督语音模型(S3M),但这些骨干网络编码了与语言相关的结构,这些结构会混淆伪造线索。我们通过语言正交化来解决这一混淆问题,这是一种无目标岭映射,用于移除S3M中投影到连续语言识别(LID)嵌入上的变异。在六种语言、六个S3M骨干网络以及所有Leave-N-Out设置中,该方法一致地降低了未见语言上的等错误率(EER)。跨语言EER与LID空间距离相关,其中正交化对更远距离的迁移带来更大的增益。
英文摘要
Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.
发表机构
- University of Southern California(南加州大学)
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。