arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

套娃语言模型套件

Matryoshka Language Model Suites

Nathan Godey, Yoav Artzi

arXiv 2608.09703首次发表:更新:

发表机构

Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Matryoshka训练框架,将尺寸递增的子模型堆叠为端到端嵌套架构,训练含5亿、15亿、30亿参数子模型的套件,其性能与基线相当,训练计算量减少36%,推测解码吞吐量提升14-26%。

AI 中文摘要

训练语言模型套件的传统方式是分别训练每个模型并独立部署。我们通过将尺寸递增的子模型堆叠为单个端到端训练的嵌套架构,提升了训练和推理效率。该Matryoshka训练框架减少了套件的总参数数量,支持在每个训练步骤中从最大模型向所有较小子模型进行低成本知识蒸馏,且因草稿模型包含在校验器内,非常适合推测解码。我们通过训练包含5亿、15亿和30亿参数子模型的Matryoshka套件验证了方法的有效性。该套件在基准性能、验证困惑度和域外困惑度上与独立训练的基线相当,同时减少了36%的训练计算量,并将推测解码的吞吐量提升了14%至26%。我们还对关键架构选择进行了 ablation 实验,为构建高性能Matryoshka语言模型套件提供了指导。

英文摘要

Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end. This Matryoshka training framework reduces the total parameter count of the suite, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding as the draft model is contained within the verifier. We validate our approach by training a Matryoshka suite comprising 500M, 1.5B, and 3B sub-models. Our suite is on par with independently trained baselines on benchmark performance and validation and out-of-domain perplexities, while using 36% less training compute and improving the throughput of speculative decoding by 14-26%. We also ablate key architectural choices, offering guidance for building strong Matryoshka LM suites.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑