发表机构
University of Sydney(悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出反向传播输出动量(BOM),将优化器动量从参数空间重定位到任务空间,在模型输出处存储预测误差的紧凑移动平均并每步重投影,减少49.7%-99.8%的优化器状态,提升微调性能1.42个百分点。
AI 中文摘要
优化器动量通常以参数大小的过去梯度移动平均值形式存储,这使得历史记录成本高昂,并将每个过去的信号固定在其计算时的坐标中。我们引入了反向传播输出动量(BOM),该方法在模型输出处存储预测误差的紧凑移动平均值,并在每一步通过当前网络重新投影该历史。批级分析刻画了这种重定位所保留和省略的信息,而实现则保留了当前的监督梯度,并可替换多种自适应优化器的第一矩分量。作为基于动量的优化器(包括那些已经压缩其状态的优化器)的插件,BOM在三种组合中将参数形状的优化器状态减少了49.7%-99.8%,并且在三个语言骨干网络上平均配对步时间减少了4.0%。它还在语言和视觉微调中提高了平均验证性能,在主要五项任务比较中提高了1.42个百分点。语言和视觉预训练研究,以及匹配的机制控制,进一步测试了该构造在不同输出空间和模型规模上的表现。
英文摘要
Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in the coordinates in which it was computed. We introduce Backpropagated Output Momentum (BOM), which instead stores a compact moving average of prediction errors at the model output and reprojects that history through the current network at every step. A batch-level analysis characterizes the information retained and omitted by this relocation, while the implementation preserves the current supervised gradient and can replace the first-moment component of several adaptive optimizers. As a plug-in for momentum-based optimizers, including ones that already compress their state, BOM reduces parameter-shaped optimizer state by 49.7-99.8% in three compositions and, averaged over three language backbones, paired step time by 4.0%. It also improves mean validation performance across language and vision fine-tuning, by 1.42 points in the primary five-task comparison. Language and vision pretraining studies, together with matched mechanism controls, further test the construction across output spaces and model scales.
Comments53 pages, 10 figures