发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出nGPT的实用训练方案,含Logit梯度预条件等技术,在14B参数的混合MoE模型上训练时,仅需约一半训练token即可达到与非归一化模型相同的验证损失,方案具可扩展性。
AI 中文摘要
归一化Transformer(nGPT)通过将模型参数向量和激活向量约束在单位超球面上实现超球表示学习。本文描述了nGPT的实用训练方案,并在现代混合Mamba-2-Transformer混合专家(MoE)模型上进行评估。该方案引入了Logit梯度预条件、对数学习率衰减、GatedAdamW、角度更新控制及可选探索机制。与采用AdamW训练的相同混合MoE架构的非归一化模型相比,总参数14B的nGPT模型使用约一半的训练token即可达到相同的验证损失。该方案在含最多14B总参数的所评估模型上具备可扩展性。
英文摘要
The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere. In this paper, we describe a practical training recipe for nGPT and evaluate it on modern hybrid Mamba-2--Transformer Mixture-of-Experts (MoE) models. The recipe introduces Logit Gradient Preconditioning, Logarithmic Learning Rate Decay, GatedAdamW, angular update control, and optional exploration mechanisms. Compared with an unnormalized model of the same hybrid MoE architecture trained with AdamW, the 30B-total-parameter nGPT model reaches the same validation loss using approximately half as many training tokens. The recipe scales across the models considered, which contain up to 30B total parameters.