arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向通用型大语言模型的深度伪造语音检测的文本声学接地

Textual Acoustic Grounding for Generalizable LLM-Based Deepfake Voice Detection

Yassine El Kheir, Xin Wang, Wanqing Ge, Tim Polzehl, Sebastian Moeller, Junichi Yamagishi

arXiv 2608.30622首次发表:更新:

发表机构

German Research Center for Artificial Intelligence (DFKI); National Institute of Informatics; Technical University of Berlin(德国人工智能研究中心; 情报信息学研究所; 柏林工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对深度伪造语音检测的域泛化问题,提出跨模态提示策略,将openSMILE提取的声学特征作为文本标记注入冻结Qwen LLM,在ITW、MLAAD基准上获16.2%宏观F1绝对提升,性能最优。

AI 中文摘要

深度伪造语音检测在未见过的领域间泛化能力较差。尽管音频大语言模型(ALLM)展现出潜力,但捕捉深度伪造检测所需的细微声学细节的连续音频嵌入,与大语言模型(LLM)语义空间之间的模态差距仍是一个关键且未被充分探索的瓶颈。我们通过将多种音频编码器与Qwen LLM(参数规模为0.5B至7B)集成进行基准测试来解决该问题。首先,我们证明仅对LLM进行微调存在域外过拟合风险,这使得冻结LLM成为更优、资源效率更高的基准。其次,为明确弥合模态差距,我们引入跨模态提示策略,通过openSMILE注入语言知识驱动的声学特征作为结构化文本标记。这种显式文本接地不仅增强了冻结基准,还使LLM微调更有效。最终,我们的方法在域外ITW和MLAAD基准上展现出最先进的鲁棒性,相较于现有ALLM基准,宏观F1值实现了超过16.2%的绝对提升,同时保持了有竞争力的域内性能。本研究中报告的所有模型均已公开可用。

英文摘要

Deepfake voice detection suffers from poor generalization across unseen domains. While Audio Large Language Models (ALLMs) show promise, the modality gap between continuous audio embeddings which capture the subtle acoustic details necessary for deepfake detection and the semantic space of LLMs remains a critical, underexplored bottleneck. We address this by benchmarking diverse audio encoders integrated with Qwen LLMs (0.5B to 7B parameters). First, we demonstrate that fine-tuning the LLM alone risks out-of-domain overfitting, making a frozen LLM a stronger, resource-efficient baseline. Second, to explicitly bridge the modality gap, we introduce a cross-modal prompting strategy that injects linguistic-knowledge-driven acoustic features (via openSMILE) as structured text tokens. This explicit textual grounding not only enhances the frozen baseline but also makes LLM fine-tuning more effective. Ultimately, our approach demonstrates state-of-the-art resilience on the out-of-domain ITW and MLAAD benchmarks, yielding over \textbf{16.2\%} absolute improvement in Macro-F1 over existing ALLM baselines while maintaining competitive in-domain performance. All models reported in this work are \href{https://huggingface.co/01Yassine/AudioLLM-Deepfake-Detection}{publicly available}.

Comments6pages + 1ref, SLT Submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑