发表机构
College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对自监督人脸表征的后门攻击研究空白,提出FIDA攻击框架,通过特征不稳定性损失优化,实现高攻击成功率与良性效用保留,可规避扰动防御,对现实多媒体应用构成威胁。
AI 中文摘要
自监督学习(SSL)模型易受后门攻击,但其在人脸表征中带来的系统性风险却鲜有研究。自监督人脸学习中身份特征的纠缠特性,为攻击的隐蔽性带来了独特挑战。为填补这一空白,本文提出FIDA(Feature Instability-Driven Attack,特征不稳定性驱动攻击),一种新型后门攻击框架。FIDA采用细微语义触发器进行注入,其核心创新是名为特征不稳定性损失(Feature Instability Loss)的新型目标函数:训练编码器,使其在攻击优化过程中,对沿扰动方向采样的触发特征提升敏感性。通过避免后门呈现出以往攻击中典型的刚性特征模式,FIDA可有效规避经评估的基于扰动的防御方法。实验表明,FIDA在所有评估设置下均实现了高攻击成功率,且通常能保留良性效用,对依赖人脸分析的现实世界多媒体应用构成重大威胁。
英文摘要
Self-supervised learning (SSL) models are vulnerable to backdoor attacks. However, the systemic risks they pose in face representation have received little attention. The entanglement of identity features in self-supervised face learning presents unique challenges for attack stealthiness. To address this gap, we propose FIDA (Feature Instability-Driven Attack), a novel backdoor attack framework. FIDA uses subtle semantic triggers for injection, but its key innovation is a novel objective called Feature Instability Loss. It trains the encoder to increase the sensitivity of triggered features along perturbation directions sampled during attack optimization . By preventing the backdoor from exhibiting the rigid feature patterns typical of previous attacks, FIDA effectively evades the evaluated perturbation-based defenses. Experiments show that FIDA achieves a high attack success rate and generally preserves benign utility across the evaluated settings , posing a significant threat to real-world multimedia applications relying on facial analysis.