arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于低秩Householder展开的无输入梯度对抗训练

Adversarial Training Without Input Gradients via Low-Rank Householder Expansions

Tiana C. Johnson, Donsub Rim

arXiv 2608.26963首次发表:更新:

发表机构

Washington University in St. Louis(圣路易斯华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于低秩Householder展开的无输入梯度对抗训练方法,消除了对输入求导及内部最大化步骤,开销低且在小ℓ²预算下可匹配多步PGD对抗训练效果。

AI 中文摘要

本研究针对训练好的深度神经网络因固有输入不稳定性产生的小范数对抗样本展开。这类样本的相对ℓ²范数较小,因此位于模型近似线性作用的输入邻域内,此时扰动仍不可察觉。我们首先证明,可通过一种名为低秩Householder展开(LRHE)的线性化方法,直接从训练好的网络参数计算这类样本,无需输入梯度迭代。该展开描述的是复合仿射映射而非单个层,其识别的方向来自前向传播中已有的激活模式。随后,我们基于此构造提出一种简单的对抗训练方案:全程不进行对输入的求导,训练仅需额外的前向评估,权重参数通过标准反向传播更新,且完全消除了常规最小-最大公式中的内部最大化步骤。我们的核心发现是存在这样的正则化项:所有放弃内部搜索的方法都通过对输入求导获取局部几何信息,而我们证明这并非必要。该正则化项每轮训练的开销相当于2.8步PGD,相较于MNIST上40步对抗训练减少了8.7倍,低于3步训练的开销。所得模型在相对ℓ²预算ε≤0.02时可匹配3步PGD对抗训练的效果,在ε≤0.012时可匹配40步训练的效果,超出该范围后性能下降,与展开的局部性一致。

英文摘要

This work concerns adversarial training against the small-norm adversarial examples that arise from the inherent input instability of a trained deep neural network. Examples in this class are small as measured in the relative $\ell^2$-norm, and therefore lie in the neighborhood of the input on which the model acts approximately linearly, the regime in which the perturbation remains imperceptible. We first show that such examples can be computed directly from the trained network parameters, without input gradient iterations, by means of a linearization called the low-rank Householder expansion (LRHE). The expansion describes the composed affine map rather than any individual layer, and the directions it identifies are read from the activation pattern already available in the forward pass. We then propose a simple adversarial training scheme built on this construction. No differentiation with respect to the input is performed at any point: training requires only additional forward evaluations, with weight parameters updated by the standard backward pass, and the inner maximization of the usual min-max formulation is eliminated entirely. That such a regularizer exists is our main finding: the methods that dispense with the inner search all obtain their local geometry by differentiating with respect to the input, and we show this is not necessary. The regularizer costs the equivalent of $2.8$ PGD steps per epoch, an $8.7\times$ reduction relative to 40-step adversarial training on MNIST and below the cost of 3-step training. The resulting models match three-step PGD adversarial training for relative $\ell^2$ budgets $\varepsilon \le 0.02$ and 40-step training for $\varepsilon \le 0.012$, falling away beyond, consistent with the locality of the expansion.

Comments25 pages, 7 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑