arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12735cs.LG

是什么让表征先验起作用?特征族、无标签不变性以及学习过程中的关键窗口

What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking

Gunner Levi Howe

首次发表
浏览论文内容

中文总结 AI 辅助

研究刻画使表征先验起作用的因素,包括特征族、无标签不变性、时间窗口等方面。通过多种实验设置,发现特征族对齐影响泛化,无标签不变性可加速,早期短暂应用先验效果最佳,为理解和改进先验方法提供依据。

中文摘要 AI 辅助

相关研究表明,学习延迟实际上是形成任务结构化表征的时间,可通过对比先验来注入。在此,我们通过188次新实验,从四个方面来刻画这种先验起作用的因素。内容方面:由错误特征族(幅度带)构建的连贯、可学习先验会阻碍泛化,就像随机划分一样,证实了先验在电路特征层面起作用的预测。监督方面:完全无标签的不变性先验——仅正样本为交换对\((a,b)\sim(b,a)\)——在15/15次实验中实现泛化,加速中位数为\(2.7\)倍,比有标签监督的先验更可靠。时间方面:先验仅在早期需要,仅在前2000个epoch(占预算的4%)应用时,能在10/10次实验中实现\(2.7\)倍加速。设置方面:解离实验在模块化乘法、不同深度和归一化变体上重复,钳位扫描量化了相关核心观点。结论是:特征族对齐决定先验是否允许泛化;不变性内容足以在无标签情况下加速;短暂的早期窗口能捕获几乎所有益处。

英文摘要

Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior. Here we characterize what makes such a prior work, across four axes, in 188 new runs. Content: a coherent, learnable prior built from the wrong feature family (magnitude bands) blocks generalization like a random partition (1/15 vs 0/20 grok; $p=0.43$ between them), confirming the companion's prediction that priors act at the level of the circuit's features. Supervision: a fully label-free invariance prior -- positives are commuted pairs $(a,b)\sim(b,a)$ only -- generalizes in 15/15 runs at a median $2.7\times$ speedup, more reliably than the label-supervised prior itself ($p=0.038$), and combined with a weight-norm clamp yields the strongest method we test (median $17\times$, 5/5) -- strongest meaning reliably fast: plain cross-entropy with a clamp matches this speed only at the exact critical norm, while the prior keeps it fast across the entire clamp range. Timing: the prior is only needed early -- applied solely during the first 2000 epochs (4% of budget) it generalizes 10/10 at $2.7\times$, beating continuous application (8/10, $1.25\times$) and a duration-matched later window ($2.1\times$). Setting: the dissociation replicates on modular multiplication and across depths and normalization variants, and a clamp sweep quantifies the companion's central claim: structure injection flattens the weight-norm delay-law exponent about 17-fold (plain cross-entropy slows $31\times$ per +10 norm units, a lower bound as higher cells are censored, versus $1.22\times$ with the prior). Honest boundary: tasks that generalize before memorizing have no delay to control. Feature-family alignment decides whether a prior permits generalization; invariance content suffices for acceleration without labels; a brief early window captures nearly all of the benefit.

↑