发表机构
University of Cambridge; University of British Columbia; Google(剑桥大学; 不列颠哥伦比亚大学; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多计算预算服务需求,提出伸缩式语言模型(TLM),通过随机前缀监督训练嵌套容量Transformer,单次训练即可在所有深度有效,质量-预算曲线面积较固定出口套件降低43-44%,GPU成本约低12%,证明训练目标而非嵌套结构决定模型弹性。
AI 中文摘要
一个已部署的语言模型通常必须服务于多种计算预算,然而为每种预算提供服务仍然意味着针对每个点进行单独的训练或压缩运行。我们训练了一个伸缩式语言模型(TLM)来成为那个连续体:一个由随机前缀监督和完整锚点监督的嵌套容量Transformer。在每一步中,容量轴的一个随机截断前缀与完整的下一词元目标一起被训练,同时还有一次全容量前向-反向传播,因此训练得到的产物在每个深度上都是一个有效的语言模型。每一步只需两次前向-反向传播,无需架构更改,推理时无需额外开销。固定出口套件,如Matryoshka语言模型套件(MLMS),占据了这个设计空间中的一个点,而这个点是有代价的:仅监督少数固定出口会使嵌套模型在其他所有地方处于随机水平(在我们的基线中困惑度为10^2-10^5)。在一个200M代理套件(20B FineWeb-Edu词元,所有方法使用相同的数据流)上,单次TLM运行在其二十个层前缀中的每一个上都是一个有效的语言模型,无论是在困惑度还是在感知困惑度的下游任务上,相对于固定出口套件,质量-预算曲线下的面积减少了43-44%,同时在全容量下与它们匹配,每次运行的GPU成本约低12%。前缀采样密度是一个调节旋钮:将其集中在少数深度上可以在那里恢复固定出口质量,但以连续体为代价,因此操作点成为训练时的选择,而非架构上的选择。这些结果表明,训练目标本身,而非嵌套结构,才是使模型具有弹性的关键。
英文摘要
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.
Comments12 pages, 4 figures, 2 tables. Code: https://github.com/ZhilinGuo/telescopic-language-models