arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19067cs.SD

什么影响伪造音频检测的性能?在持续学习设置中分析因素

What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning Setting

Yixuan Xiao, Ngoc Thang Vu

AI总结:

本研究在持续学习环境下分析攻击者架构、训练数据集、说话人多样性和任务顺序对伪造音频检测性能的影响,发现数据集伪影和任务顺序等因素显著影响系统鲁棒性。

AI中文摘要:

深度伪造音频生成技术的日益复杂,使得开发能够随时间适应的鲁棒伪造音频检测系统变得重要。本研究考察了在持续学习设置中各种因素如何影响检测系统的性能。我们关注的因素包括攻击者架构、攻击者的训练数据集、说话人多样性以及任务顺序。我们评估了使用四种不同策略训练的三个检测模型的性能,这些策略包括直接微调、单类分类、随机重放以及无遗忘学习。结果表明,伪造音频中的伪影可能源于攻击者的训练数据集,而仅仅改变攻击者架构并不足以挑战检测系统。此外,任务顺序和说话人多样性可以显著影响性能,且不同检测模型和训练策略的敏感程度各不相同。这些见解强调了在开发鲁棒检测系统时需仔细考虑这些因素的必要性。

英文摘要:

The increasing sophistication of deepfake audio generation technologies makes it important to develop robust fake audio detection systems that can adapt over time. This study examines how various factors impact the performance of detection systems in a continual learning setting. We focus on factors such as attacker architectures, attackers' training datasets, speaker diversity, and task order. We evaluate the performance of three detection models trained with four different strategies, including direct fine-tuning, one-class classification, random replay, and Learning without Forgetting. Results show that artifacts from the fake audios might arise from the attackers' training datasets, and simply changing attacker architectures does not sufficiently challenge detection systems. Moreover, task order and speaker diversity can significantly influence performance, with varying degrees of sensitivity across different detection models and training strategies. These insights underline the need for careful consideration of these factors when developing robust detection systems.

补充信息

↑