arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14753cs.SDcs.AI

用于防欺骗语音识别的大型音频语言模型

Large Audio Language Models for Spoofing-Aware Speaker Verification

Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir, Artem Dvirniak, Dmitrii Korzh, Oleg Y. Rogov

首次发表
浏览论文内容

中文总结 AI 辅助

研究用于防欺骗语音识别的大型音频语言模型,通过零样本提示等多种方式进行系统评估,发现预训练模型零样本时不佳,特定任务适应可改善,还找到多种实现竞争力性能的途径,为统一防欺骗语音识别提供基础。

中文摘要 AI 辅助

文本转语音和语音克隆的进展使高质量欺骗变得廉价且可扩展,威胁语音认证系统。现有防御主要通过深度伪造检测或防欺骗语音识别的二元对策。大型音频语言模型在相关音频任务中有潜力,但用于防欺骗语音识别未被探索。本文系统评估大型音频语言模型在零样本提示、监督适应、推理导向训练和强化学习优化下用于防欺骗语音识别的情况。结果表明预训练模型在零样本设置下接近随机水平,特定任务适应可缩小差距,还发现可通过多种途径实现有竞争力的性能。这些发现表明大型音频语言模型是统一防欺骗语音识别的有前途且可审计的基础。

英文摘要

Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.

发表机构

  • Applied AI Institute(应用人工智能研究所)
  • MIRAI(未来人工智能研究机构)
  • HSE(高等经济学院)
  • MTUCI(莫斯科国立通信与信息技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑