arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35440cs.LGcs.DC

SOLO:利用共享输出局部学习预训练十亿参数语言模型

SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning

Bojian Yin, Shurong Wang, Yuqi Pan, Guoqi Li

首次发表
浏览论文内容

中文总结 AI 辅助

SOLO通过共享最终模块的只读读出层替代私有读出层,实现无更新锁定的局部学习,在十亿参数语言模型预训练中接近反向传播性能,并显著提升内存效率和吞吐量。

中文摘要 AI 辅助

大型语言模型通过反向传播进行训练,反向传播的全局梯度协调所有层,但迫使每一层保留其激活值并等待梯度穿过所有更深层。传统的局部学习通过让每个模块使用自己的读出层来预测目标,从而消除了这种更新锁定,但尚未扩展到十亿参数的预训练。我们识别出这些私有读出层是一个关键弱点,因为它们使每个模块无法获得来自更深层模块的信息。我们提出了共享输出局部学习(SOLO),它用最终模块读出层的一个共享的、只读的副本替代这些私有读出层,该副本是唯一在整网络输出上训练的。该副本取自上一步,从最终模块传递信息,而无需在模块之间传递梯度或重新引入更新锁定。SOLO在340M到2B参数的Transformer上接近反向传播的性能,这些模型在15B个token上进行了预训练,平均零样本准确率差距在一个百分点以内,且困惑度差距随规模增大而缩小。读出层消融实验将SOLO相对于私有读出层的改进归因于共享。在没有更新锁定的情况下,每个p个流水线阶段持有O(1)个微批次的激活值,而不是O(p)个。释放的内存允许更大的微批次,其吞吐量最高可达同一分区上流水线反向传播最佳测得吞吐量的1.44倍。据我们所知,SOLO是第一种在十亿参数语言模型预训练中展示如此内存和吞吐量收益的局部学习方法。因此,局部学习成为大规模预训练中反向传播的一种实用替代方案。

英文摘要

Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update locking by training each module to predict the target through its own readout, but has not scaled to billion-parameter pretraining. We identify these private readouts as a key weakness, since they leave each module without information from deeper modules. We propose Shared-Output LOcal learning (SOLO), which replaces them with a shared, read-only copy of the final module's readout, the only one trained on the output of the whole network. Taken from the previous step, the copy transmits information from the final module without passing gradients between modules or reintroducing update locking. SOLO approaches backpropagation on Transformers of 340M to 2B parameters pretrained on 15B tokens, staying within one point in average zero-shot accuracy with a perplexity gap that narrows with scale. Readout ablations attribute SOLO's improvement over private readouts to sharing. Without update locking, each of p pipeline stages holds activations for O(1) micro-batches instead of O(p). The freed memory permits larger micro-batches, which reach up to 1.44x the best measured throughput of pipeline backpropagation on the same partition. To our knowledge, SOLO is the first local learning method to show such memory and throughput gains in billion-parameter language-model pretraining. Local learning thus becomes a practical alternative to backpropagation for large-scale pretraining.

发表机构

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑