框架化对手:一种结构感知的攻击方法
Frame the adversary: a structure-aware attack methodology
- Paris Dauphine - PSL University(巴黎多菲纳-PSL大学)
- CentraleSupélec - Paris-Saclay University(中央理工-巴黎萨克雷大学)
- Inria - Ecole Normale Supérieure - PSL University(法国国家信息与自动化研究所-巴黎高等师范学校-PSL大学)
- EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种基于结构化非正交变换和加权$\u2113_2$投影的优化框架,用于生成高效且可迁移的频率对抗攻击,并作为理论基线。\n
AI中文摘要:
基于频率的对抗性攻击近来因利用跨神经架构共享的频谱敏感性而变得流行。与空间扰动不同,基于频率的攻击暴露了更深层的脆弱性,使其对安全关键和安全性敏感应用的鲁棒性评估特别有价值。然而,现有方法通常不是作为显式捕获变换域结构的优化问题的解而推导出来的。在本文中,我们提出了一种通过专用优化框架来构建原则性频率对抗攻击的方法。我们方法的一个基石是引入一个扰动约束集,该集合与高度结构化的非正交变换相关联,这些变换以其灵活、非预定义的频率处理而闻名。我们证明了攻击表现为该集合上的加权$\u2113_2$投影,从而产生一种通用且受控的攻击生成机制。由此,我们提供了清晰的几何攻击特征,确保优化目标与扰动约束之间的对齐。我们在标准化数据集上评估了我们的框架,针对预训练模型和对抗鲁棒模型。结果表明,我们的攻击作为结构化扰动集上优化问题的解,高度有效,甚至在不同未见过的架构上也是如此。我们的方法可以作为设计和分析基于变换的攻击的理论基线,针对基本模型脆弱性,而非鲁棒性文献中通常研究的仅限架构特定的人工产物。
英文摘要:
Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly captures transform-domain structure. In this paper, we propose a methodology for crafting principled frequency-based adversarial attacks, via a dedicated optimization framework. A cornerstone of our method hinges on the introduction of a perturbation constraint set, tied to highly structured non-orthogonal transforms, well-known for their flexible, non-predefined frequency handling. We prove that the attacks emerge as weighted $\ell_2$-projections onto this set, yielding a general and controlled attack generation mechanism. By this, we provide a clear geometric attack characterization, ensuring alignment between the optimization objective and the perturbation constraint. We assess our framework on standardized datasets, for pretrained and adversarially robust models. Results highlight that our attacks, being solutions to an optimization problem, over a structured perturbation set, are highly effective, even across different, unseen architectures. Our methodology could serve as a theoretical baseline for designing and analyzing transformed-based attacks, targeting fundamental model vulnerabilities, instead of mere architecture-specific artifacts typically studied in the robustness literature.