发表机构
School of Cyber Science and Engineering, Wuhan University; Tongji university(武汉大学网络科学与工程学院; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究语音深度伪造检测在实际部署中因声学前端处理管道失真导致鲁棒性不足的问题,提出时频一致性学习框架,通过注意力驱动软对齐机制和频域结构一致性约束,提升检测模型在实际场景中的鲁棒性。
AI 中文摘要
最近,语音深度伪造检测(SDD)取得了显著进展。然而,其鲁棒性评估主要局限于可控的加性噪声场景,缺乏对实际部署中声学前端(AFE)处理管道引入的复杂失真的系统研究。本文模拟了一个统一的AFE管道,包括声学回声消除、噪声抑制、自动增益控制和语音活动检测(VAD),并对当前最先进的模型进行了全面评估。结果表明,AFE引入的非线性和时频耦合失真显著降低了检测性能。为了解决这个问题,我们提出了一个时频一致性学习(TFCL)框架,旨在学习在AFE处理前后保持稳定的不变伪造表示。我们观察到,AFE不仅引入了时间错位(例如,由VAD引起的段级偏移),还削弱或扭曲了关键的频域线索。为此,TFCL采用了一种注意力驱动的软对齐机制来捕获跨时间依赖性,同时结合频域结构一致性约束来增强特征不变性。结果,该模型能够在时间扰动和频谱失真下保持稳定的表示。广泛的实验结果表明,所提出的方法有效地减轻了AFE处理导致的性能下降,显著提高了SDD在实际场景中的鲁棒性。代码可在这个https URL上获取。
英文摘要
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lacking systematic investigation of the complex distortions introduced by acoustic front-end (AFE) processing pipelines in real-world deployments. In this work, we simulate a unified AFE pipeline comprising acoustic echo cancellation, noise suppression, automatic gain control, and voice activity detection (VAD), and conduct a comprehensive evaluation of current state-of-the-art models. The results show that the nonlinear and time-frequency coupled distortions introduced by AFE significantly degrade detection performance. To address this issue, we propose a Time-Frequency Consistency Learning (TFCL) framework, which aims to learn invariant spoofing representations that remain stable before and after AFE processing. We observe that AFE not only introduces temporal misalignment (e.g., segment-level shifts caused by VAD), but also weakens or distorts critical frequency-domain cues. To this end, TFCL employs an attention-driven soft alignment mechanism to capture cross-temporal dependencies, along with frequency-domain structural consistency constraints to enforce feature invariance. As a result, the model is able to maintain stable representations under both temporal perturbations and spectral distortions. Extensive experimental results demonstrate that the proposed method effectively mitigates the performance degradation caused by AFE processing, significantly improving the robustness of SDD in real-world scenarios. The code is available at https://github.com/JunXue-tech/TFCL.
CommentsAccepted by ACM MM 2026