AI 中文总结
研究通过向单层变压器学习模块化加法中注入不同结构内容的表征先验,因果性测试顿悟延迟是否衡量形成任务结构化表征的时间,发现只有真实结构能加速顿悟,证明顿悟延迟是形成正确表征结构的时间。
AI 中文摘要
顿悟(在训练集插值很久后才出现的泛化)可通过与结构无关的干预加速。我们通过向单层变压器学习模块化加法中注入不同结构内容的表征先验来因果性测试这一点。结果表明只有真实结构能加速顿悟,且加速与剂量有关,顿悟延迟是形成正确表征结构的时间。
英文摘要
Grokking -- generalization long after training-set interpolation -- has been accelerated by structure-agnostic interventions (gradient filtering, weight-norm clamping, geometric penalties). Whether the delay specifically measures the time to form task-structured representations has remained observational. We test it causally by injecting representational priors of varying content into a one-layer transformer learning modular addition, via a supervised-contrastive loss whose positives encode (i) the task's true structure ($(a+b) \bmod p$), (ii) a coherent-but-wrong sibling ($(a-b) \bmod p$), or (iii) a random partition -- all with identical loss form, strength, class sizes, and geometry. Whether generalization occurs follows a clean gradation: true 22/30 runs, sibling (same periodic features, wrong combination) 14/15, random (only memorizable) 0/20 (Fisher $p=1.3\times10^{-7}$). A weight-norm-matched control replaying the norm trajectory onto plain cross-entropy generalizes 0/15, ruling out the norm as mediator. Probes show structure formation precedes and predicts generalization in all runs. Only the true structure also accelerates grokking (up to $2.75\times$), but this is dose-dependent and bimodal. We then confirm the mechanism by prediction: because the acceleration is gated by a weight-norm side-effect, clamping the norm during training yields a reliable, standalone accelerator with a median $8.6\times$ speedup (up to $22\times$ on the fastest seeds, under 1000 epochs), growing monotonically as the norm is held lower; the residual stalls also vanish, though significant only pooled over the two mitigations run at both strengths ($0/40$ vs $6/20$, $p=7.7\times10^{-4}$), not per method. The grokking delay is, causally, the time to form the right representational structure -- decided at the level of features, not labels.
Commentsv3: Code and artifacts: https://doi.org/10.5281/zenodo.21132757