反向状态策略是学习算法的一部分
Backward-State Policy Is Part of the Learning Algorithm
- Beijing Tongming Lake Information Technology Application Innovation Center (TLAIC)(北京通明湖信息技术应用创新中心)
- Fudan University Institute of Systems for Advanced Computing(复旦大学先进计算系统研究所)
- Harbin Institute of Technology(哈尔滨工业大学)
- Key Lab of HCST (PKU), MOE(北京大学高可信软件技术教育部重点实验室;北京大学计算机学院)
- SCS, Peking University(上海开放计算系统研究所)
- Shanghai Institute of Systems for Open Computing
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文论证反向传播中张量读取策略(反向状态策略)是学习算法的一部分,通过理论推导和实验证明复制精度与最终损失无法判断其正确性,需逐次使用检查。
AI中文摘要:
低精度训练会舍入反向传播再次读取的张量,通常用于多个梯度;每次使用可能读取前向传播的舍入值、原始值或新的随机舍入值。这种反向状态策略看似是一个内存和精度细节,由复制精度和最终损失决定。我们认为它是学习算法的一部分,而这两项检查都无法表明其是否正确。复制精度并不能决定结果:在三对390M参数的运行中,使用模拟的FP8反向传播,当注意力机制的反向传播重用前向传播的舍入输出时训练失败,而使用来自同一分布的新舍入值时训练成功。即使是最精确的复制(即原始值本身),也可能根据我们的参考标准(即实际运行的前向传播的梯度,且梯度在通过舍入时保持不变)是错误的。例如,以低精度存储的归一化输出会馈送两个梯度:增益的梯度需要原始值,但下一层的权重梯度需要该层所乘的舍入值。最终损失(另一项检查)并不能排除对两者都读取原始值的错误:在使用这种存储方式训练的模型中,该错误持续存在,而计划的损失比较保持在预先确定的余量内。因此,我们根据这一参考标准推导出每次使用应读取哪个值,或者在前向传播固定时,哪个替代值能在平均意义上给出相同的梯度,并在单个运算符上检查这些逐次使用的要求,而无需训练。在三个使用PyTorch和Transformer Engine的测试中,这些要求预先预测了重用是否会改变反向传播相对于独立副本平均计算的内容,且每个预测都成立。因此,反向状态策略是学习算法的一部分:应逐次使用地指定和检查,而不是由复制精度和最终损失来决定。
英文摘要:
Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a memory and precision detail, settled by copy accuracy and final loss. We argue that it is part of the learning algorithm, and that neither check shows whether it is right. Copy accuracy does not decide the outcome: in three pairs of 390M runs with an emulated FP8 backward, training fails when attention's backward reuses the forward's rounded output and succeeds with a new rounding from the same distribution. Even the most accurate copy, the original itself, can be wrong by our reference: the gradient of the forward pass as it actually ran, with gradients passed through rounding unchanged. For example, a normalization output stored in low precision feeds two gradients: the gain's gradient needs the original, but the next layer's weight gradient needs the rounded value that layer multiplied. Final loss, the other check, does not rule out the error of reading the original for both: it persists in models trained with such a store, while planned loss comparisons stay within a margin fixed in advance. We therefore derive from this reference which value each use must read, or which substitute gives the same gradient on average with the forward held fixed, and check these per-use requirements on single operators, without training. In three tests using PyTorch and Transformer Engine, the requirements predicted beforehand whether reuse changes what the backward computes on average relative to an independent copy, and every prediction held. Backward-state policy is thus part of the learning algorithm: it should be specified and checked use by use, not settled by copy accuracy and final loss.