arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10287cs.LGcs.NE

训练轨迹决定可退火软先验Transformer中的电路可移除性

Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers

  • Guangdong Police College(广东警官学院)

机构由 AI 辅助整理,请以论文原文为准。

Zonglin Yang, Ziming Zhao, Wei Tang, Xunyu Jiang, Yihong Liu, Tailin Chen, Zifu Yu, Jiayu Liu

AI总结:

本研究通过可退火软先验Transformer发现,训练轨迹(而非仅最终架构)决定检索电路在移除先验后的可移除性,平滑衰减训练优于强制归零等方法。

AI中文摘要:

软位置先验可以帮助小型Transformer学习检索电路,但尚不清楚一旦移除先验,所得电路是否仍保持功能。我们使用一种可退火的软先验Transformer对此进行测试,其注意力偏置可在训练和评估期间被学习、衰减或归零。在关联回忆任务上,未强制模型在先验激活时表现良好(0.772 ± 0.020),但在门控归零时性能崩溃(0.095 ± 0.009)。平滑衰减至零的训练保持了较高的零门控准确率(0.734 ± 0.028),而强制归零训练、硬切换以及事后延续训练均无法恢复相同效果。该模式在马尔可夫归纳任务上同样出现。线性回归上下文学习提供了一个边界案例,因为零门控训练可以直接学习该任务。机制追踪显示,电路整合发生在门控达到零之后,尽管负责的头在不同随机种子间有所变化。这些结果表明,在小型离散检索任务中,电路的可移除性取决于训练轨迹,而不仅仅是最终架构。

英文摘要:

Soft positional priors can help small Transformers learn retrieval circuits, but it is unclear whether the resulting circuits remain functional once the prior is removed. We test this with an annealable soft-prior Transformer whose attention biases can be learned, faded, or zeroed during training and evaluation. On associative recall, unforced models perform well with the prior active ($0.772 \pm 0.020$) but collapse at zero gate ($0.095 \pm 0.009$). Smooth fade-to-zero training preserves high zero-gate accuracy ($0.734 \pm 0.028$), whereas forced-zero training, hard switching, and post hoc continuation fail to recover the same effect. The pattern also appears on Markov induction. Linear regression ICL provides a boundary case because zero-gate training can learn that task directly. Mechanistic traces show that circuit consolidation occurs after the gate reaches zero, even though the responsible heads vary across seeds. These results suggest that circuit removability in small discrete retrieval tasks depends on the training trajectory, not just the final architecture.

补充信息

↑