检测器学习到错误的东西:针对物理可实现攻击的抗捷径对抗训练
Detectors Learn the Wrong Thing: Shortcut-Resistant Adversarial Training Against Physically Realizable Attacks
- School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)
- State Key Lab of Intelligent Transportation System(智能交通系统国家重点实验室)
- Zhongguancun Laboratory(中关村实验室)
- School of Systems Science and Engineering, Sun Yat-Sen University(中山大学系统科学与工程学院)
- Shandong Sino-Aisa Tire Proving Ground Co.,Ltd.(山东中亚洲轮胎试验场有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对物理可实现攻击,提出实例级对比对抗训练框架InsCAT,通过SICA、ROPO和Guard等方法防止检测器将对抗纹理用作独立决策线索,经实验验证其在多场景和检测器上效果良好,提升检测可靠性。
AI中文摘要:
人工智能视觉感知系统越来越多地应用于智能交通基础设施和自动驾驶相关应用中。然而,物理可实现的对抗性外观给这些安全关键系统带来了重大的可靠性挑战。对抗训练虽有效,但对抗纹理与正人实例之间的反复共现会使检测器将纹理本身视为物体存在的证据,形成补丁纹理捷径。我们提出了InsCAT,一个实例级对比对抗训练框架,可防止检测器将对抗纹理用作独立决策线索。SICA将对抗人物特征与匹配的干净特征对齐,并将它们与仅纹理的负样本分开,而ROPO和Guard保持在线攻击压力并协调训练。我们在渲染的nuScenes、INRIAPerson、印刷服装和三个检测器系列上评估了八种独立生成的攻击纹理。InsCAT在渲染的nuScenes上实现了82.3%的平均攻击AP,超过最强基线11.1个百分点,纹理FPR从46.9%降至7.3%。物理测试的F1分数为96.6%,FPR为1.8%。跨单独训练的检测器的一致收益证明了其在具有直接推理的架构中的适用性。研究结果表明,强大的物理检测依赖于保留与目标相关的证据,同时防止对抗纹理成为独立的决策线索。
英文摘要:
AI-enabled visual perception systems are increasingly deployed in intelligent transportation infrastructure and autonomous vehicle related applications. However, physically realizable adversarial appearances pose a significant reliability challenge for these safety-critical systems. Adversarial training is effective, but repeated co-occurrence between adversarial texture and positive person instances can cause detectors to treat the texture itself as evidence of object presence, forming a patch texture shortcut. The detector may then treat texture as evidence for the target, causing false detections on texture-only inputs and weakening cross attack generalisation. We propose InsCAT, an instance-level contrastive adversarial training framework that prevents detectors from using adversarial texture as an independent decision cue. SICA aligns adversarial person features with matched clean features and separates them from texture-only negatives, while ROPO and Guard maintain online attack pressure and coordinate training. We evaluate eight independently generated attack textures on rendered nuScenes, INRIAPerson, printed garments, and three detector families. InsCAT achieves an average attack AP of 82.3% on rendered nuScenes, exceeding the strongest baseline by 11.1 points.Relative to AT-Mix, texture FPR decreases from 46.9% to 7.3%. Physical tests yield an F1 score of 96.6% and an FPR of 1.8%. Consistent gains across separately trained detectors demonstrate applicability across architectures with direct inference. The findings show that robust physical detection depends on preserving target related evidence while preventing adversarial texture from becoming an independent decision cu