发表机构
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究低精度训练中权重更新低于半ULP时坐标冻结问题,通过高精度轨迹和目标尾数长度预测冻结,以小型GPT、GPT - 2变压器等为例展示结果,还表明随机舍入可消除冻结,此条件在多种模型适用。
AI 中文摘要
在降低的浮点精度下训练可能会无声地停止学习:当梯度下降权重更新低于权重的最后一位单位(ULP)的一半时,它会舍入掉,并且该坐标会冻结,而其梯度仍不为零。这种冻结是确定性的,由每个坐标的半ULP条件控制,并且仅从高精度轨迹和目标尾数长度就可以预测,无需低精度数据。在一个按照标准AdamW加余弦方法训练、权重存储为bf16等效值的小型GPT中,训练正常进行,然后在运行到一半多一点时永久冻结,在先验预测的四步之内。在一个1.24亿参数的GPT - 2变压器中,每次优化器步骤后权重被限制在8位浮点网格上,没有主权重,两种fp8格式的密集权重在初始化时就冻结——从fp32参考中先验预测——并且验证损失平稳,而全精度不断提高。随机舍入消除了持续的冻结,同样的参考也能预测到这一点。该条件在冻结特征回归、精度跨度为128倍的尾数截断模拟器、小型网络以及MNIST上的CNN中都适用:这是低精度训练的一个可计算轴,而非扩散噪声。
英文摘要
Direct low-precision write-back can erase nonzero optimizer proposals. We ask what a high-precision reference trace establishes before a low-precision run. The exact target-code event is auditable coordinatewise on a realized target trajectory; pre-run aggregate projection also assumes the reference remains a useful counterfactual. In a controlled two-layer grid, 55/72 cells have measured and predicted post-initialization crossings: times span $384\times$, 52/55 are within 15\%, and 4/72 differ in category. Matched decoder experiments show stochastic rather than nearest write-back recovers most of the loss gap. A prospective analytic-grid E4M3 audit reuses one fp32 trace across three unseen NeoX-style seeds. It passes absolute-accuracy and skill gates (macro RMSE 0.00858) but fails directional specificity. In a target-outcome-blind comparison, a historical template has lower descriptive RMSE (0.00360) than the predeclared source predictor (0.00438); a post-outcome decomposition assigns 99.65\% of variation to common time, while a privileged matched-reference correction reaches 0.00283. Persistent-native Study~1 pairs three seeds across two schedules. Five cells are canonical; a manual sixth lacks canonical process identity, so the registered result remains inconclusive. A retrospective protocol-deviation analysis is negative because the complete constant-mid cohort is disjoint from the recovered cosine-restart cell. Study~2 reports mean full-SR/dead-zone-SR recoveries of 0.9766/0.9777 and a ratio of 1.0012, a policy contrast rather than causal mediation. Simulated-INT3 Study~3 replays six checkpoints and observes a 7.3071-nat (69.71\%) validation-loss reduction in one fixed seed. Exact events and write-back effects are auditable, but aggregate forecasts can reflect shared time rather than source-specific transfer.