arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10666cs.LGmath.OCstat.ML

解释对抗训练神经网络的显著性图稀疏性

Explaining the Saliency Map Sparsity of Adversarially-Trained Neural Networks

Yannick Lunk, Atell Yehor Krasnopolsky, Damien Garreau, Leon Bungert

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对两层ReLU网络,从理论和实验两方面解释了采用ℓ∞攻击的对抗训练神经网络的梯度显著性图稀疏性,证明其极小化子收敛到特定贝叶斯分类器,且梯度范数各向异性导致稀疏性。

中文摘要 AI 辅助

理解深度神经网络为何做出特定预测对其安全部署至关重要。在计算机视觉中,显著性图(突出显示对预测最具影响力的图像区域)仍是广泛使用的解释形式。一个经验观察是,对抗训练神经网络的梯度显著性图存在明显的稀疏性。本文针对两层ReLU网络,提出了该现象的理论解释。我们基于已有的等价关系:对抗训练等价于在经验风险最小化的基础上,加入权重衰减惩罚项和额外的对抗总变差项(该关系对某些损失函数有效)。随着数据点和神经元数量增长,且正则化参数以适当速率趋于零时,我们证明极小化子收敛到具有最小梯度范数和Barron范数的贝叶斯分类器。稀疏性的出现是因为,对于采用ℓ∞攻击的对抗训练,梯度范数是各向异性的,且更倾向于轴对齐/稀疏梯度。我们通过评估自然训练模型与对抗训练模型的梯度ℓ1范数及阈值化稀疏性,在实验中验证了上述理论发现。

英文摘要

Understanding why deep neural networks make a given prediction is of great importance for their safe deployment. In computer vision, saliency maps, which highlight the image region most influential for a prediction, remain a widely-used form of explanation. An empirical observation is the apparent sparsity of gradient saliency maps of adversarially-trained neural networks. In this paper, we propose a theoretical explanation of this phenomenon for two-layer ReLU networks. We build on the established equivalence of adversarial training to the minimization of the empirical risk with weight-decay penalization and an added adversarial total variation term -- valid for certain loss functions. As the number of data points and neurons grows and the regularization parameters are sent to zero at appropriate rates, we prove that minimizers converge to a Bayes classifier with minimal gradient and Barron norm. Sparsity appears since for adversarial training with $\ell_\infty$-attacks the gradient norm is anisotropic and favors axis-aligned / sparse gradients. We illustrate our theoretical findings experimentally by evaluating the gradient $\ell_1$-norm and thresholded sparsity of naturally versus adversarially trained models.

发表机构

  • Institute of Mathematics(数学研究所)
  • University of Würzburg(维尔茨堡大学)
  • Technical University of Munich(慕尼黑工业大学)
  • Institute of Computer Science(计算机科学研究所)
  • CAIDAS
  • Institute of Mathematics, CAIDAS(CAIDAS数学研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑