AI 中文总结
发现嵌入表和LM头的梯度几何特性,提出轻量优化器Ember,在微调、强化学习和预训练中提升帕累托前沿,仅需KB级优化器状态。
AI 中文摘要
语言模型学习离散符号上的连续程序,其中嵌入表和LM头充当它们之间的读/写接口。我们表明,该接口具有与密集隐藏权重不同的梯度几何,可以利用该几何来改进监督微调、强化学习和预训练的帕累托前沿,同时仅利用千字节的优化器状态。我们引入了Ember,一种用于嵌入和LM头矩阵的轻量级优化器,它使用O(V + D)的VRAM,而不是Adam的O(2VD),并且无需分片两个token表优化器状态。我们提供经验证据表明Ember在批大小和参数数量上有效扩展。我们表明,token的优化轨迹可以很好地由简单的一维射线描述,这与神经网络参数在高度非凸景观中导航的流行观点相反。我们提供了关于足以进行Transformer训练的优化器出奇狭窄的空间的原则性观点。最后,我们开源了分布式Ember实现,该实现与现有的ZeRO/FSDP设置干净地合并,以支持进一步的研究,网址为https://this URL。
英文摘要
Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface between them. We show that this interface has gradient geometry distinct from dense hidden weights which can be exploited to improve the Pareto frontier across supervised finetuning, RL, and pretraining, while only utilizing kilobytes of optimizer state. We introduce Ember, a lightweight optimizer for embedding and LM-head matrices that utilizes O(V + D) VRAM, instead of Adam's O(2VD), and forgoes the need to shard both token table optimizer states. We provide empirical evidence that Ember scales effectively across batch size and parameter count. We show that the optimization trajectory of tokens can be well described by a simple 1D ray, counter to the popular belief that neural net parameters navigate a heavily nonconvex landscape. We provide a principled view on the surprisingly narrow space of optimizers that suffice for Transformer training. Finally, we open-source our distributed Ember implementation that merges cleanly with existing ZeRO/FSDP setups to support further research at https://github.com/katop1234/ember