对抗前沿:用于鲁棒性评估的最小范数攻击集成
Adversarial Frontiers: Minimum-Norm Attack Ensembles for Robustness Evaluation
浏览论文内容
中文总结 AI 辅助
该研究指出对抗鲁棒性评估的现有问题,引入基于最小范数攻击综合池的统一评估框架,定义攻击前沿和防御前沿,将评估形式化为前沿逼近问题,构建攻击集成并提出防御最优性指数,在CIFAR-10和ImageNet上提供了更好的评估方法。
中文摘要 AI 辅助
对抗鲁棒性通常使用预定义的攻击集成(如AutoAttack)在单个扰动预算ε下并针对扰动范数的选择性选择进行评估。我们认为这种表述存在根本局限性。首先,鲁棒性-扰动曲线可能在不同模型间相交或以不同速率衰减,使单ε排名不稳定。其次,当前集成未提供最优性证据,与最坏情况性能存在未知差距。第三,固定攻击配置无法系统控制攻击强度与评估成本间的权衡。为解决这些局限,我们引入基于最小范数攻击综合池和跨\(\ell_0\)、\(\ell_1\)、\(\ell_2\)和\(\ell_\infty\)范数的鲁棒性-扰动曲线的统一评估框架。定义攻击前沿为攻击池对模型产生的最坏情况鲁棒性估计。将评估形式化为前沿逼近问题,构建最小范数攻击集成,在可控查询预算下逼近前沿,预算越大估计越紧。还定义防御前沿为每个扰动大小下模型集的最大鲁棒性。最后提出防御最优性指数按与防御前沿的差距对防御进行排名,无需选择参考ε。在CIFAR-10和ImageNet上,我们的集成在每个预算层级对大多数防御匹配或超过AutoAttack,以固定且可控的查询成本提供了一种查询控制、基于曲线的替代固定ε评估的方法。
英文摘要
Adversarial robustness is commonly evaluated with predefined attack ensembles, such as AutoAttack, at a single perturbation budget $\varepsilon$ and on a selective choice of perturbation norms. We argue this formulation is fundamentally limited. First, robustness--perturbation curves may intersect or decay at different rates across models, making single-$\varepsilon$ rankings unstable. Second, current ensembles provide no evidence of optimality, leaving an unknown gap to worst-case performance. Third, fixed attack configurations provide no systematic control over the trade-off between attack strength and evaluation cost. To address these limitations, we introduce a unified evaluation framework based on a comprehensive pool of minimum-norm attacks and robustness--perturbation curves across $\ell_0$, $\ell_1$, $\ell_2$ and $\ell_\infty$ norms. We define the attack frontier as the worst-case robustness estimate the attack pool produces against a model. We then formalize evaluation as a frontier-approximation problem, constructing minimum-norm attack ensembles, optimized subsets of the comprehensive pool, that approach the frontier under a controllable query budget, with larger budgets monotonically tightening the estimate. Furthermore, we define the defense frontier as the maximum robustness across the model set at each perturbation size. We finally propose the Defense Optimality Index to rank defenses by their gap to the defense frontier, providing a ranking without selecting a reference $\varepsilon$. On CIFAR-10 and ImageNet, our ensembles match or exceed AutoAttack on most defenses at every budget tier, at fixed and controllable query cost, offering practitioners a query-controlled, curve-based alternative to fixed-$\varepsilon$ evaluation.
发表机构
- Sapienza University of Rome(罗马第一大学)
- University of Cagliari(卡利亚里大学)
- University of Genoa(热那亚大学)
机构由 AI 辅助整理,请以论文原文为准。