利用梯度反转损失和多任务学习进行数据集感知的音频深度伪造检测
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
浏览论文内容
中文总结 AI 辅助
针对音频深度伪造检测中现有系统跨数据集泛化能力不足的问题,提出利用梯度反转损失和多任务学习的数据集感知框架,通过特定训练方式提升检测性能,在实验中相比基线有显著改进。
中文摘要 AI 辅助
语音合成和语音转换的进展对安全和隐私构成威胁,凸显了深度伪造检测技术的必要性。现有检测系统在单个数据集上性能强,但跨数据集泛化能力差。此前改进泛化的方法存在局限。本文提出实用的数据集感知深度伪造检测框架,仅依靠数据集标识作为监督信号进行多任务和梯度反转层训练。通过实验,与基线相比,多任务学习相对降低平均误识率13.14%,梯度反转层相对降低合并误识率5.32%,证明该方法可提高跨异构评估数据集的检测性能。
英文摘要
Recent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. Although existing detection systems achieve strong performance on individual datasets, they often fail to generalize across diverse datasets. Prior methods for improving generalization, including data augmentation, adversarial training on auxiliary factors such as language or codec types, and Mixture-of-Experts (MoE), are limited by predefined augmentation coverage, difficulties in obtaining auxiliary factors, and substantial model complexity. In this work, we propose a practical dataset-aware framework for deepfake detection. Our method targets heterogeneous datasets for which auxiliary annotations such as language, codec, or spoofing method may not be consistently available. We therefore rely only on dataset identity as a naturally available supervisory signal for multitask (MT) and gradient reversal layer (GRL) training, allowing the model to investigate both dataset-aware multitask supervision and adversarial suppression of dataset-specific information. We conduct experiments following the 2025 Speech DeepFake Arena benchmark protocol, evaluating our model across multiple evaluation datasets and reporting aggregate performance in terms of Equal Error Rate (EER), including Average EER and Pooled EER. Compared with the baseline, MT reduces Average EER by 13.14% relatively, while GRL reduces Pooled EER by 5.32% relatively. These results demonstrate that our method can improve aggregate detection performance across heterogeneous evaluation datasets, offering a practical solution for deploying reliable deepfake detection systems on diverse and unseen real-world data.