发表机构
University of California, Santa Barbara; Columbia University(加州大学圣塔芭芭拉分校; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控分解发现,在冻结SSL语音编码器中,中等多语言覆盖和掩码预测预训练优于更广更大模型,用于语音深度伪造检测。
AI 中文摘要
冻结的自监督(SSL)语音编码器是音频深度伪造检测中强大且低成本的预处理前端,最近的比较一致认为,大型、多语言、判别式编码器在域外泛化方面表现最佳。但这些比较未能同时控制编码器容量、预训练目标和多语言覆盖,只识别出哪个编码器获胜,而未隔离原因。我们提出了一种固定流程和可训练容量的受控分解。我们在四个wav2vec2系列编码器上变化多语言覆盖,参数匹配至约3.15亿。我们在两个编码器上隔离预训练目标,使用相同数据。覆盖并非单调有益,因为域外错误在约100语言规模(XLS-R)时急剧下降,但在1406语言极端(MMS)时不再改善。我们发现,中等覆盖编码器在更远的域外集上表现最强,以3.15亿参数匹配或超越5.77亿参数模型。其在远集上的领先优势在配对自助法下具有统计显著性,且在官方ASVspoof 5成本指标下在两种后端中保持。另外,在相同数据上,掩码预测预训练比对比学习泛化更好(In-the-Wild EER 26.5%对46.8%)。在此固定冻结编码器配方内,我们发现更广和更大的模型并非可靠地更好。
英文摘要
Frozen self-supervised (SSL) speech encoders are strong, low-cost front ends for audio deepfake detection, and recent comparisons agree that large, multilingual, discriminative encoders generalize best out of domain. These comparisons fail to control for encoder capacity, pretraining objective, and multilingual coverage together, identifying which encoder wins without isolating why. We present a controlled decomposition with a fixed pipeline and trainable capacity. We vary multilingual coverage on four wav2vec2-family encoders, matched to ~315M parameters. We isolate the pretraining objective on two encoders matched on identical data. Coverage does not help monotonically, as out-of-domain error drops sharply at the ~100-language scale (XLS-R) but does not improve further at the 1406-language extreme (MMS). We find that a mid-coverage encoder is strongest on farther out-of-domain sets, matching or surpassing a 577M-parameter model at 315M. Its lead on these far sets, statistically significant under paired bootstrap, and on the official ASVspoof 5 cost metric holds under two backends. Separately, masked-prediction pretraining generalizes better than contrastive on identical data (In-the-Wild EER 26.5% vs. 46.8%). Within this fixed frozen-encoder recipe, we find that broader and larger models are not reliably better.
Comments5 pages, 1 figure, 2 tables. Submitted to ICASSP 2027