arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16458eess.AScs.CL

语言正交化用于零样本跨语言音频深度伪造检测

Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

Minu Kim, Ji Sub Um, Hoirin Kim

首次发表
浏览论文内容

中文总结 AI 辅助

针对音频深度伪造检测的跨语言泛化问题,提出语言正交化方法,通过移除自监督语音模型中的语言相关变异,在多种语言和骨干网络上一致降低等错误率。

中文摘要 AI 辅助

音频深度伪造检测器需要迁移到训练中未出现的语言,因为多语言语音合成的发展速度超过了标注的反欺骗资源。尽管检测器越来越依赖自监督语音模型(S3M),但这些骨干网络编码了与语言相关的结构,这些结构会混淆伪造线索。我们通过语言正交化来解决这一混淆问题,这是一种无目标岭映射,用于移除S3M中投影到连续语言识别(LID)嵌入上的变异。在六种语言、六个S3M骨干网络以及所有Leave-N-Out设置中,该方法一致地降低了未见语言上的等错误率(EER)。跨语言EER与LID空间距离相关,其中正交化对更远距离的迁移带来更大的增益。

英文摘要

Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.

发表机构

  • University of Southern California(南加州大学)
  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑