arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GraphRectify:基于图的对抗样本检测器跨神经网络迁移方法

GraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural Networks

Arash Vashagh, Roozbeh Razavi-Far

arXiv 2610.10423首次发表:更新:

发表机构

University of New Brunswick(新不伦瑞克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GraphRectify通过图结构学习分类器中间特征表示,实现对抗样本检测器在不同骨干网络间的有效迁移,在多数情况下优于从头训练,尤其在不同架构家族间迁移时增益显著。

AI 中文摘要

对抗样本检测器通常与其训练时所用的分类器骨干网络紧密绑定,这限制了当受保护模型被替换或升级时检测器的复用。由于不同网络通常产生不兼容的内部表示,因此直接在骨干网络之间迁移此类检测器具有挑战性。我们提出GraphRectify,一种基于图的框架,用于在分类器骨干网络之间迁移对抗图像检测器。GraphRectify学习分类器中间特征的结构化表示,并将来自新骨干网络的特征表示适配到在原始模型上学习的检测器上,从而实现检测器的复用。我们在多个数据集、骨干网络架构和对抗攻击(包括联合针对分类器和检测器的检测器感知自适应攻击)上评估GraphRectify。在整个评估矩阵中,GraphRectify相比在新骨干网络上从头训练检测器以及所评估的迁移消融方法,取得了更高的总体ROC-AUC。在不同骨干网络家族之间的迁移以及数据充足的情况下,增益尤为显著。相比之下,在数据最受限的设置中,从头训练仍具有竞争力。这些结果表明,对抗检测知识可以在异构分类器架构之间有效迁移,而无需在受保护骨干网络发生变化时重新学习。

英文摘要

Adversarial example detectors are often tied to the classifier backbone they were trained on, limiting reuse when the protected model is replaced or upgraded. Directly transferring such detectors across backbones is challenging because different networks generally produce incompatible internal representations. We propose GraphRectify, a graph-based framework for transferring adversarial image detectors across classifier backbones. GraphRectify learns a structured representation of intermediate classifier features and adapts representations from a new backbone to the detector learned on the original model, enabling detector reuse. We evaluate GraphRectify across multiple datasets, backbone architectures, and adversarial attacks, including detector-aware adaptive attacks that jointly target the classifier and detector. Across the complete evaluation matrix, GraphRectify achieves higher aggregate ROC-AUC than training a detector from scratch on the new backbone and the evaluated transfer ablations. The gains are particularly strong for transfers between different backbone families and when sufficient data are available. In contrast, training from scratch remains competitive in the most data-limited settings. These results show that adversarial detection knowledge can transfer effectively across heterogeneous classifier architectures rather than being relearned whenever the protected backbone changes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑