发表机构
INESC TEC; Faculty of Engineering, University of Porto (FEUP); Unilabs(INESC TEC; 波尔图大学工程学院; Unilabs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过60个检测器的受控实验,揭示源-目标兼容性(如骨干网络、预训练、训练数据)显著影响对抗样本迁移成功率,并指出源平均会低估目标脆弱性,强调源模型选择的重要性。
AI 中文摘要
深度伪造检测器仍然容易受到基于迁移的黑盒攻击,在这种攻击中,对抗样本在源替代模型上生成,并迁移到攻击者未知的目标模型。然而,源-目标兼容性如何影响攻击成功率仍未被充分理解。先前的研究评估的检测器池有限,且很少将架构因素与训练因素分开考虑。我们对跨60个检测器的对抗迁移性进行了受控评估,这些检测器涵盖六种骨干网络、两种预训练方案和五种训练数据配置,使用两种攻击程序:AutoAttack (AA) 和带期望变换的Carlini--Wagner攻击 (CW--EOT)。匹配比较显示,当源和目标共享完全相同的骨干网络、架构家族、预训练方案或训练数据时,迁移性显著更高。这种兼容性结构依赖于攻击方式:在AA下,完全骨干网络兼容性影响最大,而在CW--EOT下,共享预训练和训练数据影响最大。当迁移性在非目标源上取平均时,AA下的平均攻击成功率 (ASR) 为7.21%,CW--EOT下为19.52%。相比之下,结合两种攻击的多源oracle在排除完全骨干网络和训练数据匹配后,平均ASR达到64.48%,表明源平均可能大幅低估目标脆弱性。我们发布了240,000张对抗扰动图像、完整的成对迁移结果、检测器配置和评估代码。这些发现确立了源-目标兼容性和源模型选择作为可信的基于迁移的黑盒鲁棒性评估的核心维度。
英文摘要
Deepfake detectors remain vulnerable to transfer-based black-box attacks, in which adversarial examples are generated on a source surrogate model and transferred to a target model, unknown to the attacker. Yet how source--target compatibility shapes attack success remains poorly understood. Prior studies evaluate limited detector pools and rarely disentangle architectural from training factors. We conduct a controlled evaluation of adversarial transferability across 60 detectors spanning six backbones, two pretraining regimes, and five training-data configurations, using two attack procedures: AutoAttack (AA) and the Carlini--Wagner attack with Expectation over Transformation (CW--EOT). Matched comparisons reveal significantly higher transfer when source and target share an exact backbone, architecture family, pretraining regime, or training data. This compatibility structure is attack-dependent: exact backbone compatibility has the largest effect under AA, whereas shared pretraining and training data have the largest effects under CW--EOT. When transfer is averaged across non-target sources, mean attack success rate (ASR) is $7.21\%$ under AA and $19.52\%$ under CW--EOT. By contrast, a multi-source oracle combining both attacks attains a \(64.48\%\) mean ASR after excluding exact backbone and training-data matches, showing that source averaging can substantially understate target vulnerability. We release 240,000 adversarially perturbed images, complete pairwise transfer results, detector configurations, and evaluation code. These findings establish source--target compatibility and source-model selection as central dimensions of credible transfer-based black-box robustness evaluation.