源定向轨迹扰动:以一阶代价实现语音深度伪造检测中的域泛化
Source-Directed Trajectory Perturbation at First-Order Cost for Domain Generalization in Speech Deepfake Detection
浏览论文内容
中文总结 AI 辅助
针对语音深度伪造检测中域偏移导致的性能下降,提出源定向轨迹扰动方法WC-MLDG-2P,以两倍梯度评估成本模拟最坏情况,分别降低平均EER 23.4%和7.6%。
中文摘要 AI 辅助
语音深度伪造检测器在测试数据分布与训练数据分布不一致时,往往会出现准确率下降的问题。用于域泛化的元学习(MLDG)通过使用情景化的元训练和元测试划分来模拟域偏移,显示出一定的前景。然而,MLDG的元目标仅考虑单一的干净适应轨迹,并未考虑其端点周围的局部损失景观。为解决这一局限性,我们首先提出了一种显式的最坏情况MLDG变体,称为WC-MLDG-4P,该变体同时扰动元训练和元测试状态,但每个情景需要四次梯度评估。随后,我们引入了WC-MLDG-2P,这是对显式鲁棒目标的一种源定向替代方案。它将干净的MLDG端点向局部更高源损失状态移动,并在该处评估元测试梯度,从而保留了一阶MLDG的两次评估成本。通过一阶展开,我们将所得的梯度变化与沿源梯度的元测试方向曲率联系起来,而无需显式计算Hessian矩阵。相对于MLDG,WC-MLDG-2P在使用XLSR-AASIST和XLSR-Conformer-TCM时,分别实现了相对平均EER降低23.4%和7.6%。
英文摘要
Speech deepfake detectors often lose accuracy when the distribution of the test data differs from that of the training data. Meta-learning for domain generalization (MLDG) shows promise by simulating domain shifts with episodic meta-train and meta-test splits. However, the MLDG meta-objective considers only a single clean adaptation trajectory and does not account for the local loss landscape around its endpoints. To address this limitation, we first propose an explicit worst-case MLDG variant, dubbed WC-MLDG-4P, which perturbs both the meta-train and meta-test states but requires four gradient evaluations per episode. We then introduce WC-MLDG-2P, a source-directed alternative to the explicit robust objective. It shifts the clean MLDG endpoint toward a locally higher source-loss state and evaluates the meta-test gradient there, retaining the two-evaluation cost of the first-order MLDG. A first-order expansion relates the resulting gradient change to meta-test directional curvature along the source gradient without explicitly computing a Hessian. Relative to MLDG, WC-MLDG-2P achieves relative mean-EER reductions of 23.4% with XLSR-AASIST and 7.6% with XLSR-Conformer-TCM, respectively.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。