发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨预训练中注意力softmax的量化近似,通过K区间注意力分析校准与代理位置的影响,发现固定窗口校准和归一化后代理可最小化验证损失差距。
AI 中文摘要
低精度Transformer系统越来越多地对注意力矩阵乘法进行量化,而softmax通常仍保持较高精度。在预训练期间,近似softmax会改变训练模型的梯度及其前向计算。我们通过K区间注意力研究这种交互,该注意力使用K+1个网格值来近似指数函数。我们改变每行网格校准、插值与硬舍入,以及直通代理相对于归一化的位置。我们推导相应的反向规则,包括校准导数,并在模型、数据和优化器匹配的预训练实验中比较这些选择。分离行极值使前向计算保持不变,但会导致验证损失的延迟增加。在K=4时采用硬舍入,最小-最大校准和归一化前代理会产生较大的损失差距;改变其中任一选择可大幅减少该差距。在124M参数和2.5B训练令牌下,固定窗口校准配合归一化后代理在K=4时相对于softmax产生+0.019 nats的验证损失差距,而配合归一化前代理在K=16时产生+0.004 nats的差距。
英文摘要
Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interval attention, which approximates the exponential using K+1 grid values. We vary per-row grid calibration, interpolation versus hard rounding, and the placement of a straight-through surrogate relative to normalization. We derive the corresponding backward rules, including calibration derivatives, and compare these choices in pretraining experiments matched on model, data, and optimizer. Detaching the row extrema leaves the forward computation unchanged but produces a delayed increase in validation loss. With hard rounding at K=4, min-max calibration and a pre-normalization surrogate incur a large loss gap; changing either choice substantially reduces it. At 124M parameters and 2.5B training tokens, fixed-window calibration with a post-normalization surrogate yields a validation loss gap of +0.019 nats relative to softmax at K=4, and with a pre-normalization surrogate yields +0.004 nats at K=16.