向过去学习:一种带有自蒸馏正则化的代理引导对抗防御框架
Learn from the Past: A Proxy Guided Adversarial Defense Framework with Self Distillation Regularization
- School of Software Technology, Dalian University of Technology(大连理工大学软件学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有对抗训练存在的训练不稳定、灾难性过拟合问题,提出名为LAST的代理引导对抗防御框架,利用目标模型历史状态作代理,结合自蒸馏正则化,显著提升模型鲁棒性与训练稳定性,适配多种场景。
AI中文摘要:
对抗训练(Adversarial Training, AT)是增强深度学习模型鲁棒性的关键技术,已在实际应用中被广泛采用。然而,现有主流AT方法依赖对目标模型的防御进行直接迭代更新,常面临训练不稳定、灾难性过拟合等问题。\n在此背景下,本研究发掘了利用目标模型历史状态作为代理以提供有效初始化和防御先验的潜力,由此构建了通用的代理引导防御框架LAST(Learn from the Past,向过去学习)。具体而言,LAST将代理模型的响应作为动态学习的快速权重,持续修正目标模型的更新方向。此外,我们提出了一种自蒸馏正则化防御目标,该目标经过巧妙设计,无需借助外部教师模型即可引导代理模型的更新轨迹,从而缓解灾难性过拟合对性能的影响。\n大量实验与消融研究表明,该框架可显著提升模型鲁棒性(例如在CIFAR10和CIFAR100数据集上,鲁棒准确率分别最高提升9.2%和20.3%)与训练稳定性。这些提升在不同模型架构、更大规模数据集、扰动强度及攻击模式下均一致存在,证实了LAST可同时优化单步与多步AT策略的能力。代码将发布于https://github.com/callous-youth/LAST。
英文摘要:
Adversarial Training (AT), pivotal in fortifying the robustness of deep learning models, is extensively adopted in practical applications. However, prevailing AT methods, relying on direct iterative updates for target model's defense, frequently encounter obstacles such as unstable training and catastrophic overfitting. In this context, our work illuminates the potential of leveraging the target model's historical states as a proxy to provide effective initialization and defense prior, which results in a general proxy guided defense framework, `LAST' ({\bf L}earn from the P{\bf ast}). Specifically, LAST derives response of the proxy model as dynamically learned fast weights, which continuously corrects the update direction of the target model. Besides, we introduce a self-distillation regularized defense objective, ingeniously designed to steer the proxy model's update trajectory without resorting to external teacher models, thereby ameliorating the impact of catastrophic overfitting on performance. Extensive experiments and ablation studies showcase the framework's efficacy in markedly improving model robustness (e.g., up to 9.2\% and 20.3\% enhancement in robust accuracy on CIFAR10 and CIFAR100 datasets, respectively) and training stability. These improvements are consistently observed across various model architectures, larger datasets, perturbation sizes, and attack modalities, affirming LAST's ability to consistently refine both single-step and multi-step AT strategies. The code will be available at~\url{https://github.com/callous-youth/LAST}.