延迟优化器状态传输塑造短程训练决策
Delayed Optimizer-State Transport Shapes Short-Horizon Training Decisions
浏览论文内容
中文总结 AI 辅助
该研究发现优化器状态与未来数据的延迟传输可改变Transformer短程训练决策,在Math--Code数据集上提升损失,为有限视野训练干预提供了机制依据。
中文摘要 AI 辅助
自适应优化器会在矩变量中保留梯度历史,这使得损失权重的局部变化能够改变后续的更新。本文研究这种延迟传输是否大到足以改变预期的短程决策。针对已确定的未来小批量序列,我们通过完整的模型-优化器状态对八步AdamW轨迹求导,并在独立评估前选择了暴露匹配的Math--Code损失调度。在12个未使用的0.3M Transformer历史数据中,与感知优化器的即时导数相比,完整传输在10/12个历史数据中降低了不重叠token的损失(平均收益为4.71×10⁻⁴;精确单侧符号检验,p=0.0193)。两种控制方法的作用频率相同,但在60/96个窗口中选择了不同的调度。交叉检查点-未来路径测试将这种重新排序归因于优化器状态与近期未来数据之间的相互作用,而独立的Ising--CNN实验表明,删除矩状态传输会破坏准确的响应预测。完整传输分数还会将精确回滚的获胜者集中在更大的候选库中,将有限幅度的评估聚焦于候选短名单。因此,在这些已确定的短路径上,优化器记忆与近期未来数据顺序是训练状态中可操作的组成部分,为何时需要有限视野而非单步干预提供了基于机制的标准。
英文摘要
Adaptive optimizers retain gradient history in moment variables, allowing a local change in loss weighting to alter later updates. We examine whether this delayed transport is large enough to change prospective short-horizon decisions. On committed future-minibatch sequences, we differentiate eight-step AdamW trajectories through the complete model--optimizer state and select exposure-matched Math--Code loss schedules before independent evaluation. Across 12 unused 0.3M Transformer histories, full transport lowers token-disjoint loss relative to an optimizer-aware immediate derivative in 10/12 histories (mean benefit $4.71\times10^{-4}$; exact one-sided sign test, $p=0.0193$). The two controllers act equally often but select different schedules in 60/96 windows. Crossed checkpoint--future-path tests attribute this reordering to the interaction between optimizer state and near-future data, while an independent Ising--CNN experiment shows that deleting moment-state transport destroys accurate response prediction. Full-transport scores also concentrate exact-rollout winners in larger candidate libraries, focusing finite-amplitude evaluation on a shortlist. On these committed short paths, optimizer memory and near-future data order are therefore actionable components of the training state, providing a mechanism-based criterion for when finite-horizon rather than one-step intervention is required.
发表机构
- School of Physics, Beihang University(北京航空航天大学物理学院)
机构由 AI 辅助整理,请以论文原文为准。