AI 中文总结
该研究针对对抗训练为何减少叠加的问题,基于Ilyas等人的特征分类法,通过实证解释揭示其因果链,为理解对抗样本与叠加的关系提供了新视角。
AI 中文摘要
对抗样本及其起源的研究仍是开放的研究领域。机械可解释性,尤其是叠加现象,为解决该问题提供了新途径。Gorton与Lewis(2025)证明对抗样本源于叠加,并通过实验表明对抗训练会减少叠加,但未提供为何会发生此情况的机械解释。我们基于Ilyas等人(2019)的特征分类法,提出一种实证解释,追溯以下因果链:对抗训练放弃非鲁棒特征,导致用于表示的总特征减少,进而减少了叠加。
英文摘要
The study of adversarial examples and their origins remains an open area of research. Mechanistic interpretability, and superposition in particular, offers new avenues for approaching this problem. Gorton & Lewis (2025) demonstrate that adversarial examples arise from superposition and show empirically that adversarial training reduces superposition, yet provide no mechanistic account of why this occurs. We present an empirical explanation inspired by the feature taxonomy of Ilyas et al. (2019), tracing the following chain of causalities: adversarial training abandons non-robust features, leading to fewer total features to represent, resulting in less superposition.
Comments9 pages, 4 figures. Accepted at the COLM 2026 Workshop on AI Interpretability (AIW)