伪类感知加权攻击(DAWA):打破伪类防御中的安全汇合
Dummy-Aware Weighted Attack (DAWA): Breaking the Safe Sink in Dummy Class Defenses
- Tsinghua University(清华大学)
- University of Macau(澳门大学)
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出DAWA,一种新的评估方法,通过自适应加权同时攻击真实标签和伪标签,有效突破伪类防御机制,降低其鲁棒性评估值。
AI中文摘要:
对抗鲁棒性评估面临新防御范式带来的关键挑战,现有评估方法难以有效检测这些新防御的局限性。本文揭示,基于伪类的防御方法在传统评估策略如AutoAttack下会显著高估鲁棒性。根本限制在于这些攻击仅关注误导真实标签,与防御机制完美契合——成功攻击被伪类捕获。为解决这一缺口,我们提出伪类感知加权攻击(DAWA),一种新的评估方法,通过自适应加权同时攻击真实标签和伪标签。大量实验表明,DAWA有效突破这一防御范式,将领先伪类防御在CIFAR-10下的鲁棒性从58.61%降至29.52%(l_infty扰动,epsilon=8/255)。本文提供了更可靠的基准评估该新兴防御类别的方法,并强调了鲁棒性评估方法持续演化的必要性。
英文摘要:
Adversarial robustness evaluation faces a critical challenge as new defense paradigms emerge that can exploit limitations in existing assessment methods. This paper reveals that Dummy Classes-based defenses, which introduce an additional "dummy" class as a safety sink for adversarial examples, achieve significantly overestimated robustness under conventional evaluation strategies like AutoAttack. The fundamental limitation stems from these attacks' singular focus on misleading the true class label, which aligns perfectly with the defense mechanism--successful attacks are simply captured by the dummy class. To address this gap, we propose Dummy-Aware Weighted Attack (DAWA), a novel evaluation method that simultaneously targets both the true label and dummy label with adaptive weighting during adversarial example synthesis. Extensive experiments demonstrate that DAWA effectively breaks this defense paradigm, reducing the measured robustness of a leading Dummy Classes-based defense from 58.61% to 29.52% on CIFAR-10 under l_infty perturbation (epsilon=8/255). Our work provides a more reliable benchmark for evaluating this emerging class of defenses and highlights the need for continuous evolution of robustness assessment methodologies.