arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19818cs.SDcs.AI

CoRELoop:用于音频深度伪造检测的参数高效受控循环精化

CoReLoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Amphion Technology Co., Ltd.(安飞昂科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Kunyu Feng, Yuxiang Wang, Li Wang, Wan Lin, Zhizheng Wu

AI总结:

针对音频深度伪造检测泛化难题,提出CoReLoop方法,在冻结的SSL检测器上通过轻量级循环精化模块和低秩适配器实现参数高效精化,在14个跨域测试集上将EER从4.85%降至3.74%。

AI中文摘要:

对于音频深度伪造检测器而言,泛化到未见过的攻击仍然具有挑战性,而收集覆盖所有潜在攻击的训练数据是不切实际的。我们探索了在已训练的基于SSL的检测器中,无需额外数据或改变其原始参数的情况下进行循环精化。然而,在我们的诊断中,直接重用编码器输出作为输入会降低检测性能。我们提出了CoReLoop,通过将循环输入适配到冻结的编码器、控制状态更新以及将精化输出与冻结的分类器对齐,使这种重用变得有效。通过在原始数据上仅训练轻量级精化模块和循环特定的低秩适配器,CoReLoop在保持检测器原始首轮预测的同时实现了额外的精化。在14个跨域测试集上,24层模型通过两次传递将池化等错误率(EER)从4.85%降低到3.74%,可训练参数约为598M中的10M。为了选择性地应用这种精化,一个可选的停止头为每个话语选择深度,实现了3.73%的池化EER,平均传递次数为1.18次。

英文摘要:

Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an already-trained SSL-based detector without additional data or changes to its original parameters. However, directly recycling encoder outputs as inputs degrades detection in our diagnostic. We propose CoReLoop, which makes this reuse effective by adapting recurrent inputs to the frozen encoder, controlling state updates, and aligning refined outputs with the frozen classifier. By training only lightweight refinement modules and loop-specific low-rank adapters on the original data, CoReLoop enables additional refinement while preserving the detector's original first-pass prediction. On 14 cross-domain test sets, the 24-layer model reduces pooled equal error rate (EER) from 4.85% to 3.74% with two passes, with approximately 10M trainable parameters out of 598M. To selectively apply this refinement, an optional halting head chooses the depth for each utterance, achieving 3.73% pooled EER with an average of 1.18 passes.

补充信息

↑