arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

鲁棒性优化防御中的共享漏洞:一次突破暴露整个家族

Shared Vulnerabilities in Robustness-Optimized Defenses: One Breach Exposes the Family

Hanrui Wang, Ruihao Zheng, Shuo Wang, Isao Echizen, Xingbo Dong, Zhe Jin

arXiv 2607.18339首次发表:更新:

AI 中文总结

研究对抗鲁棒性优化防御中存在的共享漏洞风险,通过引入新协议、攻击方法及映射,发现各家族内部有自然可转移性,净化防御风险严重,未来防御应兼顾漏洞多样性和转移隔离,而非仅优化个体鲁棒性。

AI 中文摘要

对抗鲁棒性优化旨在在对抗扰动下保持正确预测,通过对抗训练和对抗净化等方法取得了显著的鲁棒性提升。然而,研究发现这些提升可能会在防御中产生共享漏洞。一旦一个代表性的鲁棒性优化防御被有效突破,整个家族可能都会暴露。为研究此风险,引入了更严格的仅转移协议和简单的自适应攻击PGDTransfer,还引入了对抗敏感性映射(AdvSMs)。研究发现各鲁棒性家族内部存在自然可转移性,净化防御的风险已很严重,未来防御应将漏洞多样性和仅转移隔离作为安全目标。

英文摘要

Adversarial robustness optimization aims to preserve correct prediction under adversarial perturbations, and has produced substantial robustness gains through methods such as adversarial training and adversarial purification. However, we identify a new security risk: these gains can create shared vulnerabilities across defenses. Once one representative robustness-optimized defense is effectively breached, the broader family may become exposed. Studying this risk requires separating genuine transferability from distortion-induced degradation and from the algorithmic gains of sophisticated attacks. We therefore introduce stricter transfer-only protocols and a deliberately simple adaptive attack, PGDTransfer, to test whether robustness-optimized defenses share transfer-only vulnerability under controlled conditions. We further introduce Adversarial Sensitivity Maps (AdvSMs) to visualize and quantify shared alignment beyond differentiable classifiers, including stochastic and non-differentiable defenses. Across adversarially trained classifiers, purification-based defenses, and LVLMs with robust visual encoders, we identify natural transferability within each robustness family, i.e., transfer that arises even with simple PGD-style optimization rather than specialized transferable-attack design. The risk is already severe for purification: PGDTransfer reaches an average transfer attack success rate of $80.4\%$ across filtering-, compression-, and diffusion-based purifiers under $ε=4/255$, suggesting that purifier defenses may no longer provide reliable protection. As attacks improve, currently stronger robustness families may face the same risk. Future defenses should therefore treat vulnerability diversity and transfer-only isolation as security objectives, rather than optimizing only individual robustness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑