发表机构
StepFun(StepFun)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
KITE通过KV不变扩展实现模型从小到大的高效训练,并降低推理成本;SST双塔解码器在可比计算下优于更大基线模型。
AI 中文摘要
扩展语言模型不仅仅是最终质量的问题:架构选择决定了在训练、提示处理和自回归解码过程中花费多少计算量以达到特定的模型质量。理想的模型架构应降低上述所有计算成本,以便扩展到更大的模型,同时确保更大的模型确实优于较小的基线模型。我们引入了KV不变变压器扩展(KITE),一种实现这一目标的扩展范式。它从较小的规模训练模型到较大的规模(即通过升级回收节省训练成本),同时将新添加的参数放置在不会影响注意力KV的区域。因此,在推理过程中,预填充KV仅依赖于模型的较小部分,从而节省了推理成本。作为具体实例,我们提出了步进规模变压器(SST),一种双塔解码器,其中一个塔生成KV,另一个塔读取它们。在可比的累积训练计算下,SST是一个67B的MoE模型,每个解码令牌具有2.15B的活动主体参数,其训练损失低于分别具有1.48B和2.02B活动主体参数的47B和63B MoE变压器,同时估计推理成本分别降低6.7%和31.6%。
英文摘要
Scaling a language model is not only a question of final quality: the architectural choice determines how much computation is spent during training, prompt processing, and autoregressive decoding to achieve certain model quality. An ideal model architecture should lower all above computation costs to facilitate scaling to a larger model, while ensure the larger model indeed outperforms smaller baselines. We introduce KV-Invariant Transformer Expansion (KITE), a scaling paradigm that achieves this goal. It trains the model from a smaller size to a larger size (i.e., saving training costs via upcycling), while places newly added parameters in regions that do not affect attention KV. Consequently, during inference, prefilling KV only relies on the smaller part of the model, so the inference costs are saved. As a concrete instantiation, we present Step Scale Transformer (SST), a two-tower decoder in which one tower produces KV and the other reads them. At comparable cumulative training compute, SST, a 67B MoE model with 2.15B active body parameters per decode token, achieves lower training loss than 47B and 63B MoE Transformers with 1.48B and 2.02B active body parameters, respectively, while reducing estimated inference cost by 6.7% and 31.6%.