发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究挑战Transformer中块作用与位置绑定的传统观点,提出基于循环深度折叠的莫比乌斯学习,同一组块在不同数据流中有不同应用时机,实现深度角色叠加,实验表明其在分布式训练中有优势,开辟新设计空间。
AI 中文摘要
基于Transformer的语言模型沿有序深度轴组织计算,浅块和深块常发挥不同表征作用。本文挑战了这些作用必须与块在有序序列中的位置绑定的传统观点。引入基于循环深度折叠的训练架构莫比乌斯学习,不同数据流遵循循环移位的块顺序。同一组块对一些数据流在序列早期应用,对另一些在晚期应用,实现深度角色叠加。在四 worker 实验中,用Muon在25亿FineWeb令牌上训练修改后的GPT - 2小模型(1.24亿参数),莫比乌斯学习在更多Transformer块序列轮次时验证损失更低。这表明块组不必局限于序列中固定的浅或深角色,基于循环深度折叠开辟了新设计空间,且该结构适合内存受限的分布式训练。
英文摘要
Transformer-based language models organize computation along an ordered depth axis, where shallow and deep blocks often develop distinct representational roles. We challenge the conventional view that these roles must remain tied to a block's position in the ordered sequence. We introduce Mobius Learning, a training architecture based on cyclic depth folding, in which different data streams follow cyclically shifted block orders. The same block group is therefore applied early in the block sequence for some data streams and late for others, so it is optimized in both shallow and deep roles, a phenomenon we call depth-role superposition. Surprisingly, in four-worker experiments with a modded GPT-2 small (124M) model trained on 2.5B FineWeb tokens using Muon, Mobius Learning achieves lower validation loss than a fixed-order looped Transformer at larger numbers of Transformer block-sequence passes. This counterintuitive result shows that a block group need not remain confined to one fixed shallow or deep role within the block sequence and opens a new design space based on cyclic depth folding. Crucially, this structure makes Mobius Learning particularly well suited to memory-constrained distributed training: raw training data remain local, while each worker stores one block group rather than the complete Transformer block stack.