AI 中文总结
研究针对神经网络快速学习问题,基于Transformer注意力与热力学系统同构,将注意力对数its方差视为比热Cv,提出CvAdamW变体,经改进解决失败模式,实现快速学习加速,还提出无任务特定超参数的重新表述并评估,效果良好。
AI 中文摘要
快速学习是指神经网络在记住训练数据很久之后才出现的延迟泛化,它浪费了数千个训练轮次且难以预测。基于Transformer注意力与热力学系统形式同构的最新结果,我们将注意力对数its的方差视为比热Cv,并表明其峰值可靠地先于泛化转变。我们引入了CvAdamW,它是AdamW的变体,可在线监测Cv,并在检测到相变时通过动态缩放权重衰减来注入热能。通过严格的迭代开发过程,我们识别出三种失败模式,并通过记忆门和指数移动平均减震器解决了它们。在模运算(a+b mod 97)中,CvAdamW在4000个训练轮次的预算内于第2802个轮次实现了快速学习,而基线模型从未实现快速学习。我们进一步提出了一种无任务特定超参数的尺度不变z分数重新表述,并在10对种子上进行了评估。配对分析表明,冷启动变体将平均快速学习延迟减少了257个轮次(6.0%;中位数为166个轮次;Wilcoxon p=0.049,Cohen's d=0.68,自举95% CI [53,489]),在10个种子中有8个得到了改善;在这个单一任务中,所有10个种子的Cv在快速学习之前达到峰值。我们的结果表明,神经网络可能会暴露即将发生的泛化转变的可检测前兆,并且基于物理动机的比例干预可以在固定的计算预算内促进泛化。代码和数据是公开的。
英文摘要
Grokking -- the delayed generalization of neural networks long after they have memorized their training data -- wastes thousands of training epochs and is notoriously unpredictable. Building on the recent result that Transformer attention is formally isomorphic to a thermodynamic system, we treat the variance of attention logits as a specific heat Cv and show that its peak reliably precedes the generalization transition. We introduce CvAdamW, a drop-in AdamW variant that monitors Cv online and injects thermal energy by dynamically scaling weight decay when a phase transition is detected. Through a strictly iterative development process we identify three failure modes -- initialization noise, mini-batch micro-ripples, and slingshot blinding -- and resolve them with a memorization gate and an exponential-moving-average shock absorber. On modular arithmetic (a+b mod 97), CvAdamW enables grokking at epoch 2802 in a 4000-epoch budget where the baseline never groks. We further propose a scale-invariant z-score reformulation that removes task-specific hyperparameters, and evaluate it across 10 paired seeds. A paired analysis shows the cold-start variant reduces mean grokking latency by 257 epochs (6.0%; median 166 epochs; Wilcoxon p=0.049, Cohen's d=0.68, bootstrap 95% CI [53,489]), improving 8 of 10 seeds; on this single task Cv peaks before grokking in all 10 seeds. Our results indicate that neural networks may expose detectable precursors of impending generalization transitions, and that a physically motivated, proportional intervention can facilitate generalization within a fixed compute budget. Code and data are public.
Comments9 pages, 4 figures, 2 tables. Code and data: https://github.com/baymaxbyte/cbo_core