arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MRCert:通过类型特定掩码实现对抗性修补样本的部署后修补鲁棒性认证

MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking

Qilin Zhou, Zhengyuan Wei, Haipeng Wang, Zhuo Wang, Shuo Liu, W. K. Chan

arXiv 2610.10617首次发表:更新:

发表机构

City University of Hong Kong; The Education University of Hong Kong(香港城市大学; 香港教育大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MRCert是首个基于掩码的认证恢复防御者,通过类型特定设计实现部署后对抗性修补样本的鲁棒性认证,在ImageNet修补尺寸16像素时获35.1%对抗性认证准确率,优于SOTA的PatchCURE。

AI 中文摘要

在部署后阶段,深度学习模型的输入可能是也可能不是经过对抗性修补的。在修补边界内对这类输入进行修补鲁棒性认证,可验证其标签的良性,且应保持较高的预测准确率。然而,现有的基于平滑和基于掩码的恢复防御者无法同时实现这两点:前者会大幅降低预测准确率,后者则无法验证对抗性修补输入返回标签的良性。我们提出MRCert,这是首个基于掩码的认证恢复防御者,证明了同时实现上述两点的可行性。与所有现有工作对两类输入(良性和对抗性修补样本)应用统一条件进行认证不同,MRCert在部署后阶段为两类输入分别推断深度学习模型的类型特定必要属性,并通过标签恢复与认证函数对的新型面向类型设计,将这些属性正式关联以验证标签的良性。在不产生平滑导致的干净准确率下降的情况下,实验结果证实,MRCert在修补尺寸为16像素时,于ImageNet数据集上实现了35.1%的对抗性认证准确率,而当前最优方法PatchCURE则完全失效。

英文摘要

In post-deployment time, inputs to deep learning models may or may not be adversarially patched. Patch robustness certification on such inputs within a patch bound can verify their label benignity and should retain high prediction accuracy. However, existing smoothing-based and masking-based recovery defenders cannot achieve both simultaneously: they degrade the prediction accuracy much and cannot verify the benignity of the returned label of an adversarially patched input, respectively. We propose MRCert, the first masking-based certified recovery defender that shows the feasibility of achieving both. Unlike all existing works to apply a common condition across both types of input (benign and adversarially patched samples) for certification, MRCert infers type-specific necessary properties of deep learning models for both types in post-deployment time and formally relates them to verify the label benignity through a novel type-oriented design of label recovery and certification function pair. Without incurring the degradation in clean accuracy caused by smoothing, experimental results confirm that MRCert achieves 35.1\% adversarial certified accuracy on ImageNet at patch size 16 pixels, whereas the SOTA PatchCURE fails completely.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑