arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型预训练中的贪婪局部学习:差距与目标设计

Greedy Local Learning for Language Model Pretraining: Gaps and Objective Design

Jihwan Moon, Sheir A. Zaheer, Jinmyoung Lee, Gunhee Kim, Chan Y. Park

arXiv 2610.04867首次发表:更新:

发表机构

INFOCZ Inc.; Seoul National University; KC ML2(INFOCZ公司; 首尔大学; KC ML2)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过125M和400M参数规模的实证,揭示了贪婪局部学习在语言模型预训练中的损失差距随块数增加而扩大,并发现Transformer辅助、MTP目标限制及逐块执行可有效缩小差距、降低内存。

AI 中文摘要

贪婪的逐块局部学习将网络拆分为由局部辅助损失训练的梯度隔离块,删除块之间的反向传播:阶段间通信变为仅前向,每个块可以独立更新其优化器,这些特性与去中心化模型并行训练直接相关。局部学习在图像分类上与端到端反向传播相当,在小型Transformer上已知会以更差的最佳损失换取并行加速。这种损失差距在更大规模的自回归语言模型(LM)预训练中如何表现,以及哪些辅助设计能缩小差距,尚未被测量。我们提出了一项在125M和400M参数规模下、块数K∈{1,2,4}、采用Chinchilla最优预算的token预算匹配的实证研究,将辅助设计分解为网络架构和训练目标两个因素。我们观察到:(i)与端到端训练的差距从K=2到K=4增加了一倍以上,但在K=4时从125M到400M有所缩小;(ii)用基于Transformer的辅助替代MLP辅助是强大的网络侧干预,恢复了22-41%的差距;(iii)多token预测(MTP)辅助目标在第一个块边界有帮助,而在更深边界添加则有害,将其限制在第一个块产生了最佳的K=4配置(在400M时+0.062对比+0.075 nats);(iv)部署风格的逐块执行将激活内存减少最多2.2倍。我们将这些结果定位为经验支撑的方法方向而非最终方法:局部目标应选择性地在边界间施加未来预测压力,同时抵制绕过预测内容的捷径。

英文摘要

Greedy block-wise local learning splits a network into gradient-isolated blocks trained by local auxiliary losses, deleting the backward pass between blocks: inter-stage communication becomes forward-only and every block can step its optimizer independently, properties directly relevant to decentralized model-parallel training. Local learning is competitive with end-to-end backpropagation on image classification, and on small Transformers it is known to trade a worse best loss for parallel speedup. How this loss gap behaves in autoregressive language model (LM) pretraining at larger scale, and which auxiliary designs reduce it, has not been measured. We present a token-budget-matched empirical study at 125M and 400M parameters with $K \in \{1,2,4\}$ blocks at Chinchilla-optimal budgets, factorizing the auxiliary design into network architecture and training objective. We observe: (i) the gap to end-to-end training more than doubles from $K=2$ to $K=4$, but at $K=4$ shrinks from 125M to 400M; (ii) replacing an MLP auxiliary with a Transformer-based one is a strong network-side intervention, recovering 22-41% of the gap; (iii) a multi-token-prediction (MTP) auxiliary objective helps at the first block boundary, whereas adding it at deeper boundaries hurts, and restricting it to the first block yields the best $K=4$ configuration ($+0.062$ vs. $+0.075$ nats at 400M); and (iv) deployment-style per-block execution reduces activation memory by up to $2.2\times$. We frame these results as an empirically grounded method direction rather than a finalized method: local objectives should apply future-predictive pressure selectively across boundaries while resisting shortcuts that bypass predictive content.

CommentsAccepted at the NeurIPS 2026 Workshop on Collaborative, Open, and Decentralized Training of Foundation Models (CODEC-FM). 13 pages, 3 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑