arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过双教师的信息瓶颈蒸馏提升对抗攻击下的鲁棒性/准确率权衡

Improving the Robustness/Accuracy Tradeoff Against Adversarial Attacks Using Information Bottleneck Distillation Through Dual Teachers

Vincent Ryusuke Takahashi, Yoshinari Takeishi, Jun'ichi Takeuchi, Kave Salamatian

arXiv 2607.27737首次发表:更新:

发表机构

Joint Graduate School of Mathematics for Innovation; Kyushu University; Faculty of Information Science and Electrical Engineering; Université Savoie Mont Blanc(联合创新数学研究生院; 九州大学; 信息科学与电气工程学院; 萨瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过引入干净教师扩展信息瓶颈蒸馏框架,在CIFAR-10/100数据集上提升了干净样本分类准确率,保持对抗样本准确率,与现有双教师蒸馏方法竞争力相当。

AI 中文摘要

深度神经网络(DNNs)在经典机器学习问题中取得了显著成功,但已知其易受对抗攻击影响。文献中提出的对策,特别是Kuang等人引入的信息瓶颈蒸馏(IBD),在提升对抗输入鲁棒性的同时,会降低干净输入的分类准确率。本研究扩展了IBD框架,在由对抗训练得到的鲁棒教师模型的蒸馏过程中,引入仅用干净输入训练的额外教师模型(干净教师);通过跨层注意力矩阵,将干净教师与鲁棒教师的特征迁移至学生模型。在CIFAR-10和CIFAR-100数据集上的实验结果显示,与原始IBD相比,所提方法提升了干净样本的分类准确率,同时保持了对抗样本的类似准确率;此外,所提方法在干净准确率与鲁棒准确率的调和均值方面,与包括近期双教师蒸馏框架B-MTARD在内的最先进方法具有竞争力;还分析了不同训练设置对注意力模块的不同影响。

英文摘要

Deep neural networks (DNNs) have achieved remarkable success in classical machine learning problems. However, they are known to be vulnerable to adversarial attacks. Countermeasures proposed in the literature, notably Information Bottleneck Distillation (IBD) introduced by Kuang et al., degrade the classification accuracy on clean inputs while improving the robustness to adversarial inputs. In this work, we extend the IBD framework by introducing an extra teacher model (clean teacher) trained with only clean inputs, into the distillation process from a robust teacher model trained by adversarial training. The features of both clean and robust teachers are transferred to the student through a cross-layer attention matrix. Experimental results on the CIFAR-10 and CIFAR-100 datasets show that the proposed method improves classification accuracy on clean samples compared to the original IBD, while maintaining similar accuracy on adversarial samples. Furthermore, our methods are competitive with state-of-the-art approaches, including the recent dual-teacher distillation framework B-MTARD, particularly in terms of the harmonic mean between clean and robust accuracy. We also analyze the impact of different training settings that have different influences on the attention module.

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑