AI 中文总结
该研究通过对Fashion-MNIST上训练的卷积神经网络开展PGD攻击的轨迹级分析,发现轨迹级诊断指标无法独立衡量对抗鲁棒性,仅失效步数分布可清晰区分鲁棒性层级,轨迹分析应作为标准鲁棒性测量的补充工具。
AI 中文摘要
投影梯度下降(PGD)被广泛用于评估对抗鲁棒性,通常通过最终对抗准确率来衡量,但这一指标无法捕捉攻击全过程中模型的行为。近期研究提出了轨迹级诊断指标,如损失演化、梯度对齐和失效步数,以深入洞察对抗优化动力学。然而,这些指标是否能可靠指示鲁棒性强度仍不明确。我们对在Fashion-MNIST上训练的卷积神经网络开展PGD攻击的轨迹级研究,在多个鲁棒性 regime 中对比了干净训练和对抗训练模型:采用严格的20步PGD评估(含随机初始化和多次重启以测量鲁棒性,单次初始化以记录轨迹),为每个模型记录3000个干净正确样本的完整PGD轨迹,分析攻击迭代中的损失演化、梯度对齐和失效时机。结果显示模型间存在清晰的鲁棒性层级,但轨迹指标对其识别的贡献并不均等:具有显著不同鲁棒准确率的对抗训练模型,其平均损失轨迹和梯度对齐模式在数量上相似;相反,失效步数分布能更清晰地区分鲁棒性 regime,直接反映对抗扰动的功能抗性。这些发现表明,轨迹级诊断指标描述的是优化几何,而非独立衡量对抗鲁棒性,其可解释性取决于鲁棒性 regime、攻击强度和多指标评估;轨迹级分析应作为补充诊断工具,结合上下文解读,而非替代标准鲁棒性测量。
英文摘要
Projected Gradient Descent (PGD) is widely used to evaluate adversarial robustness, typically via final adversarial accuracy, which does not capture model behaviour throughout the attack. Recent work proposes trajectory-level diagnostics, such as loss evolution, gradient alignment, and steps-to-failure, for deeper insight into adversarial optimisation dynamics. However, whether these diagnostics reliably indicate robustness strength remains unclear. We conduct a trajectory-level investigation of PGD attacks on convolutional neural networks trained on Fashion-MNIST. We compare clean-trained and adversarially-trained models across multiple robustness regimes, using rigorous 20-step PGD evaluations with random initialisation and multiple restarts for robustness measurement, and single-initialisation trajectory recording for diagnostics. We record full PGD trajectories across 3000 clean-correct samples per model and analyse loss evolution, gradient alignment, and failure timing across attack iterations. Our results reveal a clear robustness hierarchy across models; however, trajectory metrics do not contribute equally to its identification. Mean loss trajectories and gradient alignment patterns appear quantitatively similar across adversarially-trained models with substantially different robust accuracies. In contrast, steps-to-failure distributions provide a clearer separation of robustness regimes, directly reflecting functional resistance to adversarial perturbation. These findings indicate that trajectory-level diagnostics describe optimisation geometry but do not independently measure adversarial robustness. Their interpretability depends on robustness regime, attack strength, and multi-metric evaluation. Trajectory-level analysis should be a complementary diagnostic tool, interpreted in context, rather than a replacement for standard robustness measurements.
Comments16 pages, 3 figures