arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Grokked 幻觉:真实均衡缓解灾难性遗忘

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting

Xiaotian Zhang, Lai Shun Chan, Yue Shang, Entao Yang, Ge Zhang

arXiv 2607.29503首次发表:更新:

AI 中文总结

该研究以模运算的 grokking 为场景,发现相同饱和性能下,高熵模型比 AdamW 训练的 Transformer 更能抵御灾难性遗忘,其注意力与 MLP 层有效秩更高,揭示完美泛化不代表同等鲁棒性。

AI 中文摘要

神经网络通常通过训练和测试性能进行评估,但这些指标无法揭示学习到的表示的鲁棒性。近期研究表明,以玻尔兹曼熵量化的参数空间中占据更大体积的解,相比传统优化得到的解往往表现出更优的泛化性,这一现象被称为高熵优势。本研究探究该优势是否超越泛化性,具体考察模型的鲁棒性——即模型在后续训练以获取新信息时保留所学知识的能力。我们以模运算中的 grokking 为受控场景,设计噪声注入实验,评估 AdamW 训练的 Transformer 与从 Wang-Landau 分子动力学采样的、具有相同饱和性能的高熵模型之间的鲁棒性差异。通过强制两个模型用随机标签完全记忆新数据,我们发现 AdamW 训练的模型遭受灾难性遗忘,原任务测试准确率从 100% 降至 75% 以下,而高熵模型保持约 95% 的测试准确率。我们将这种表面泛化背后的隐藏脆弱性称为“grokked 幻觉”。通过对神经网络权重进行奇异值分解,我们发现高熵神经网络在注意力层和 MLP 层中,无论是否注入噪声,都具有显著更高的有效秩,表明更丰富的特征表示可作为抵御灾难性遗忘的缓冲。我们的发现表明,完美的泛化性并不意味着同等的鲁棒性,为训练模型抵御干扰的鲁棒性提供了新视角。

英文摘要

While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown that solutions occupying larger volumes in parameter space, as quantified by Boltzmann entropy, often exhibit superior generalizability compared to those reached by conventional optimization, a phenomenon known as the high entropy advantage. Here we ask whether this advantage persists beyond generalization. Specifically, we investigate models' robustness, the ability to retain the learned knowledge when the model is subsequently trained to acquire new information. Using grokking in modular arithmetic as a controlled setting, we design a noise injection experiment to evaluate the robustness difference between AdamW-trained transformers and high-entropy model sampled from Wang-Landau Molecular Dynamics with identical saturated performance. By forcing both models to fully remember new data with random labels, we find that AdamW-trained models suffer from catastrophic forgetting, with original task test accuracy dropping from 100% to below 75%, whereas the high-entropy models maintain approximately 95% test accuracy. We term this hidden fragility behind apparent generalization the "grokked illusion." Through singular value decomposition of the neural network weights, we discover that high-entropy neural networks possess significantly higher effective rank in attention and MLP layers both before and after noise injection, indicating richer feature representations can serve as a buffer against catastrophic forgetting. Our findings demonstrate that perfect generalization does not imply equal robustness, offering a new perspective on what makes a trained model robust to interference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑