Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model
训练混合:将小规模支架式预训练运行重组为更大的语言模型
机构 * Google(谷歌公司) ; School of Computing, Dublin City University(都柏林城市大学计算机学院)
AI总结 本研究提出训练混合(MoT)框架,将Transformer划分为层块在冻结对齐器内独立训练后重组,在13亿参数Gemma模型上验证其可重组为可用语言模型,可复用对齐器实现计算优势,用于研究可复用训练单元。
Comments Accepted at the Workshop on Methods and Opportunities at Small Scale (MOSS), COLM 2026