发表机构
Barcelona Supercomputing Center (BSC)(巴塞罗那超级计算中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究利用自监督学习隐空间的轨迹动力学检测音频深度伪造,通过LSTM预测器与MLP结合的系统在多基准测试中达到最优性能,验证了生理约束可作为检测信号。
AI 中文摘要
人类语音的产生受生理机制约束,会在声学信号上形成特有的时间结构。我们假设这些约束会在自监督学习(SSL)模型的隐空间中表现为结构化的轨迹动力学,而合成语音会违反这些约束且可被检测到。为验证该假设,我们仅用真实语音训练因果长短期记忆(LSTM)下一帧预测器(阶段1),使用深度伪造专用的SSL骨干网络Wav2Vec2-Large-AntiDeepfake,并采用相同特征与静态全局平均池化基线对比,以分离时间建模的贡献。我们还加入了监督阶段2,即使用带标签数据在冻结的LSTM内部状态上训练多层感知机(MLP),以表征伪造监督的作用。我们的系统在六个基准测试中取得了具有竞争力或达到当前最优的性能:ASVspoof 2019/2021、Codecfake、In-the-Wild、MLAAD-EN和Deepfake-Eval-2024,其中包括在ASVspoof 2021上达到已发表的最佳等错误率(EER)0.75%,且值得注意的是,仅用真实语音训练的阶段1在DE2024上超过了同一骨干网络的已发表监督基线(30.35%)。在近域基准测试中,静态与动态方法表现相当;在具有多种合成方法的更难跨语料库基准测试中,轨迹动力学带来了显著提升,证实时间生理约束携带了超越语句级统计的检测信号。
英文摘要
Human speech production is constrained by physiology, giving rise to characteristic temporal structure on acoustic signals. We hypothesise that these constraints manifest as structured trajectory dynamics in the latent space of Self-Supervised Learning (SSL) models, and that synthetic speech violates them detectably. To test this hypothesis, we train a causal Long Short-Term Memory (LSTM) next-frame predictor on bonafide speech only (Stage 1), using the deepfake-specialised SSL backbone Wav2Vec2-Large-AntiDeepfake, and compare against a static global-average-pooling baseline using identical features, thus isolating the contribution of temporal modelling. A supervised Stage 2, which trains a Multi-Layer Perceptron on the frozen LSTM internal states using labelled data, is included to characterise the role of spoof supervision. Our system achieves competitive or state-of-the-art performance across six benchmarks: ASVspoof 2019/2021, Codecfake, In-the-Wild, MLAAD-EN, and Deepfake-Eval-2024, including best published EER on ASVspoof 2021 (0.75\%) and, notably, Stage 1 trained on bonafide speech only surpasses the published supervised baseline from the same backbone on DE2024 (30.35\%). On near-domain benchmarks, static and dynamic approaches perform comparably. On harder cross-corpus benchmarks with diverse synthesis methods, trajectory dynamics provide substantial gains, confirming that temporal physiological constraints carry detection signal beyond utterance-level statistics.
Comments5 pages, 1 figure