发表机构
Ben-Gurion University of the Negev; Afeka the Academic College of Engineering; Avignon University(内盖夫本-古里安大学; 阿费卡工程学术学院; 阿维尼翁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出 Spooftral 模型,基于 Voxtral 音频语言模型结合指令引导与 DoRA 适配,在 ASVspoof5 上达到 4.25% 等错误率,验证了 ALM 框架用于语音欺骗检测的可行性。
AI 中文摘要
近年来,自监督学习(SSL)的对抗措施(CMs)表现出强大的性能。然而,在面对未见过的欺骗攻击和条件不匹配的情况下,它们往往性能下降。本研究考察了 Voxtral 音频语言模型(ALM)框架用于欺骗检测,作为将 CM 能力整合到 ALM 框架中的一步。我们分析了 Voxtral 如何通过音频-文本处理捕获欺骗线索,并提出了一种指令引导的方法,利用标签序列似然来评估真实和欺骗语音。在 ASVspoof 数据库上的实验表明,在没有任务特定适配的情况下,LLM 层强调语义表示,与基于 Whisper 的音频编码器相比,降低了欺骗判别声学线索的可分离性。因此,经过语言模型处理后,与欺骗相关的信息变得不那么可分离。我们还对 Voxtral 模型应用了使用权重分解低秩适配(DoRA)的轻量级适配,并提出了 Spooftral 模型,在 ASVspoof5 评估集上实现了 4.25% 的等错误率(EER)。
英文摘要
Self-supervised learning (SSL) countermeasures (CMs) have shown strong performance in recent years. However, they often show degraded performance while facing unseen spoofing attacks and mismatched conditions. This study examines the Voxtral audio-language model (ALM) framework for spoofing detection, as a step toward combining CM capabilities within the ALM framework. We analyze how Voxtral captures spoofing cues through audio-text processing and propose an instruction-guided approach that uses label-sequence likelihoods to evaluate bonafide and spoofed speech. Experiments on the ASVspoof databases show that without task-specific adaptation, the LLM layers emphasize semantic representations, reducing the separability of spoof-discriminative acoustic cues compared to the Whisper-based audio encoder. Consequently, spoofing-related information becomes less separable after language-model processing. We also applied lightweight adaptation using weight-decomposed low-rank adaptation (DoRA) to the Voxtral model and propose the Spooftral model, achieving an equal error rate (EER) of 4.25% on the ASVspoof5 evaluation set.
Comments8 pages, 3 figures, 5 tables. Accepted to the Spoken Language Technology (SLT) 2026