arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MOSAIC:融合生物语音特征与自监督表示的可解释多令牌交叉注意力统一语音防欺骗方法

When and why do handcrafted cues help self-supervised anti-spoofing? A causal and faithfulness analysis

Yugwon Won

arXiv 2607.04314首次发表:更新:

发表机构

AI Security R&D Team(人工智能安全研发团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有语音防欺骗融合方法可解释性差、跨模态学习受限问题,提出MOSAIC可解释多令牌交叉注意力框架,实现跨特征对齐可视化,在多数据集上取得优异性能。

AI 中文摘要

当前语音防欺骗领域的主流趋势是将自监督学习(SSL)骨干网络(例如WavLM)与人工设计特征相融合,但这类融合方式通常在线索与层级的交互层面缺乏透明度,且简单的拼接操作会限制跨模态学习效果。本文提出MOSAIC(面向多令牌的集成交叉注意力语音防欺骗框架),这是一种可解释的多令牌交叉注意力框架,它将152维生物语音特征向量划分为6个语义组查询令牌(Praat、相位、LFCC均值/标准差、子带均值/标准差),并在13个经均值-标准差池化的WavLM-Large Transformer层级上对其执行注意力计算,将这些层级作为键/值。最终得到的6×13注意力矩阵可直观展示语音线索与模型层级的对齐关系;针对单令牌激活值的z-score分析显示,生物语音/相位令牌在真实语音上的激活程度更高,而频谱/通道令牌在欺骗语音上的激活程度更高——该方法可得到单线索、单层级的归因结果,对现有融合方法形成了有效拓展。通过联合焦点损失、双域LA/PA对抗分类器以及仅针对真实语音的VAE正则器开展训练,MOSAIC在ASVspoof 2019 LA/PA数据集上的等错误率(EER)达到1.93%/1.98%,该单一统一模型性能接近PA专用的当前最优模型(LFCC-CMR,EER为1.34%),同时在LA任务上仍保持竞争力;其在ASVspoof 2021 LA/DF/PA数据集上的EER分别为9.28%/6.21%/40.09%。

英文摘要

Most spoofing countermeasures now place a light classifier on top of a self-supervised (SSL) speech encoder. A growing line of work adds handcrafted acoustic features and fuses them by cross-attention, partly because the attention weights appear to explain which cues the model relies on. Two things about such fusion remain untested: whether the handcrafted features contribute anything once a strong SSL encoder is already in place, and whether the attention map faithfully reflects what the classifier actually uses. We study both questions with MOSAIC, a single model that handles logical access (LA) and physical access (PA) attacks by projecting a 152-dimensional biophonetic vector into six query tokens attending over thirteen intermediate layers of WavLM-Large. Rather than pursuing state-of-the-art accuracy, we treat the model itself as the object of analysis and apply inference-time interventions that zero out selected components. Removing the handcrafted branch lowers the EER by 0.67 percentage points under LA, so WavLM alone is sufficient there, but raises it by 0.76 points under PA, where the branch supplies replay channel evidence the encoder does not carry. The 6x13 attention map is highly stable under resampling (bootstrap rank correlation 0.99), yet its faithfulness varies by domain: layer attention weights predict the causal importance of each layer under PA (rho = +0.62) but not under LA (rho = -0.28). The handcrafted branch and its attention-based explanation are therefore trustworthy for replay attacks and not for synthetic ones. On standard benchmarks the model is mid-range on LA and ahead of the official baselines on cross-source deepfakes (6.21% EER on ASVspoof 2021 DF). The contribution is not a new architecture but a checking procedure: verify handcrafted cues and attention maps by intervention before trusting either.

Comments24 pages, 4 figures, 4 tables. Substantially revised from v1 (5 pages, 2 figures): reframed from an architecture proposal to a causal and faithfulness analysis; adds full ASVspoof 5 evaluation (680,774 utterances) and bootstrap confidence intervals

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑