破坏注意力:对检测 Transformer 中编码器注意力的基于规避的对抗攻击
Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers
浏览论文内容
中文总结 AI 辅助
本文提出首个针对检测 Transformer 编码器注意力的对抗攻击,在不可察觉扰动下破坏注意力结构,使 DETR-R50、DINO-Swin-L 等主流检测器 mAP 大幅下降,性能优于现有攻击。
中文摘要 AI 辅助
对抗漏洞仍是神经网络安全部署的主要隐患,尤其在目标检测领域——该任务是诸多安全关键系统的核心组成部分。检测 Transformer 已成为主流目标检测器,但其对抗鲁棒性却相对未被充分研究。现有多数攻击针对检测输出,而非这类模型独具的注意力机制。本文提出首个在不可察觉的有界 ℓ∞ 扰动下直接优化编码器注意力目标的攻击方法:它不通过可见补丁引入攻击者控制的 sink token,而是驱动模型自身注意力朝向被破坏的目标。本文认为,编码器注意力集中了模型的空间推理能力,因此破坏它比单独扰动检测输出会在检测流程中产生更具破坏性的传播效果。在相同扰动预算和迭代次数下,本文攻击将 COCO 数据集上 DETR-R50 的 mAP 从 42.1 降至 0.97,比现有最强攻击的 mAP 降低幅度约 4 倍。本文进一步表明,该漏洞并非特定于某一破坏目标:在分散、重排序、排列和峰抑制这 4 种性质不同的目标下,检测性能均降至 mAP 3 以下,说明该弱点源于注意力结构本身被破坏,而非任何单一目标。最后,本文证明该攻击可跨注意力形式泛化:将 DINO-Swin-L 的 mAP 从 56.8 降至 1.44,而现有最强攻击仅降至 7.3,在密集注意力和可变形注意力上均达到当前最优。
英文摘要
Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in many safety-critical systems. Detection transformers have emerged as leading object detectors, yet their adversarial robustness remains comparatively underexplored. Most existing attacks target the detection output rather than the attention mechanism that makes these models distinctive. In this paper, we introduce the first attack that directly optimizes an encoder-attention objective under an imperceptible, bounded $\ell_\infty$ perturbation. Rather than introducing an attacker-owned sink token through a visible patch, it drives the model's own attention toward a corrupted target. We argue that encoder attention concentrates the model's spatial reasoning, so corrupting it propagates through the detection pipeline more disruptively than perturbing the detection output alone. Our attack reduces DETR-R50 mAP on COCO from 42.1 to 0.97, a $\sim 4\times$ reduction in resulting mAP over the strongest existing attack under an identical perturbation budget and iteration count. We further show that this vulnerability is not specific to a particular corruption objective: across four qualitatively distinct targets, dispersion, re-ranking, permutation, and peak-suppression, detection consistently drops below 3 mAP, suggesting that the weakness arises from disrupting the attention structure itself rather than from any single target. Finally, we demonstrate that the attack generalizes across attention formulations, reducing DINO-Swin-L from 56.8 to 1.44 mAP against 7.3 for the strongest prior attack, establishing state-of-the-art on both dense and deformable attention.
发表机构
- Queensland University of Technology(昆士兰科技大学)
- Agency for Science, Technology and Research, Singapore(新加坡科技研究局)
- National University of Singapore(新加坡国立大学)
- Indraprastha Institute of Information Technology, Delhi(英迪拉普拉斯塔信息技术学院(德里))
机构由 AI 辅助整理,请以论文原文为准。