AI 中文总结
针对共享预算和局部约束下的对抗扰动分配问题,提出基于固定梯度先验的重要性引导分配机制,在不扩大预算下重分配扰动至敏感区域,显著提升攻击成功率。
AI 中文摘要
在共享 $\ell_1$ 预算下的对抗优化不仅需要决定使用多少扰动,还需要决定有限的预算应花费在何处。当单个输入坐标受到局部幅度约束时,这一分配问题变得尤为重要,因为局部约束限制了扰动集中在少数位置上的程度。我们引入了一种重要性引导的分配机制,该机制利用固定的干净梯度先验将扰动引导至模型敏感区域,同时保持可行扰动集不变。一个中心化分配目标鼓励在高于平均重要性的位置进行扰动,并抑制在其他位置的不必要支出,从而重新分配而非扩大可用预算。在常见的容量受限威胁设置下,跨越十个鲁棒模型-数据集配置,所提方法相较于匹配的APGD和PMA基线,攻击成功率提高了2.52到17.70个百分点。分配分析表明,这些增益伴随着高重要性区域中扰动质量的显著增加,而全局 $\ell_1$ 消耗并未增加。机制消融进一步表明,中心化非均匀重分配提供了部分收益,而模型派生的重要性则带来了额外的改进。这些结果将扰动分配确定为共享预算、局部受限威胁模型下对抗优化的一个独特且实际相关的维度。
英文摘要
Adversarial optimization under a shared $\ell_1$ budget requires deciding not only how much perturbation to use, but also where that limited budget should be spent. This allocation problem becomes particularly important when individual input coordinates are subject to local magnitude constraints, which restrict the extent to which perturbation can be concentrated on a small number of locations. We introduce an importance-guided allocation mechanism that uses a fixed clean-gradient prior to steer perturbation toward model-sensitive regions while leaving the feasible perturbation set unchanged. A centered allocation objective encourages perturbation at above-average importance locations and discourages unnecessary expenditure elsewhere, thereby redistributing rather than enlarging the available budget. Across ten robust model--dataset configurations under a common capacity-limited threat setting, the proposed method improves attack success over matched APGD- and PMA-based baselines by $2.52$ to $17.70$ percentage points. Allocation analysis shows that these gains are accompanied by substantially greater perturbation mass in high-importance regions without increased global $\ell_1$ consumption. Mechanism ablations further show that centered non-uniform redistribution provides part of the benefit, while model-derived importance yields an additional improvement. These results identify perturbation allocation as a distinct and practically relevant dimension of adversarial optimization under shared-budget, locally constrained threat models.