arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22854cs.LG

合适规模的思考:跨后训练大语言模型的摊销蒸馏

Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs

Yan Zhou, Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak, Marco Fumero, Francesco Locatello, David Alvarez-Melis

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出ADAPT框架,可单次蒸馏生成跨后训练大语言模型的多规模多变体模型,实现平滑规模插值与自适应规模选择,优化长文本推理的计算-准确率权衡。

中文摘要 AI 辅助

大语言模型(LLM)的实际部署需要一系列后训练变体——指令调优、推理调优和对话式模型——每种变体都有多种规模,以满足不同的延迟和内存预算。独立生成每个(变体,规模)对的成本过高,因此模型系列通常对每个后训练变体仅包含少数粗粒度规模。Boomerang蒸馏(Kangaslahti等人,2026)针对基础模型沿规模轴降低了这一成本,它通过模型规模插值,从单个师生对构建中间规模的模型,无需额外训练。但该方法仍将每个后训练变体视为单独的优化对象。我们提出ADAPT——跨后训练大语言模型的摊销蒸馏(Amortized Distillation Across Post-Trained LLMs)——一个在模型系列的规模和后训练变体两个维度上摊销蒸馏的框架,通过单次蒸馏生成针对K种后训练变体的L种插值规模的L×K个模型。ADAPT包含两个组件:一是两阶段蒸馏过程,通过预训练对齐和监督微调蒸馏构建后训练学生模型,使其在生成和推理任务上实现平滑的规模-性能插值;二是权重增量初始化,通过将基础模型的蒸馏诱导权重变化转移到从不同后训练变体初始化的学生模型,近似实现跨后训练变体的构建。生成的插值模型连续体还支持推理时的自适应模型规模选择,改善了长文本推理任务的计算-准确率权衡。

英文摘要

Practical deployment of large language models (LLMs) requires families of post-trained variants---instruction-tuned, reasoning-tuned, and chat-style models---each at multiple sizes to meet diverse latency and memory budgets. Producing each (variant, size) pair independently is prohibitive, so model families typically span only a handful of coarse-grained sizes per post-trained variant. Boomerang distillation (Kangaslahti et al., 2026) reduces this cost along the size axis for base models. Through model size interpolation, it constructs models of intermediate sizes from a single teacher-student pair without additional training. However, it still treats each post-trained variant as a separate object of optimization. We introduce ADAPT---Amortized Distillation Across Post-Trained LLMs---a framework for amortizing distillation across both axes of a model family: size and post-training variant, producing $L \times K$ models for $L$ interpolated sizes across $K$ post-trained variants with a single distillation run. ADAPT combines two components. First, a two-phase distillation procedure constructs post-trained students through pre-training alignment and supervised fine-tuning distillation, enabling smooth size--performance interpolation on generation and reasoning tasks. Second, weight-delta initialization approximates this construction across post-trained variants by transferring the distillation-induced weight change from the base model to students initialized from different post-trained variants. The resulting continuum of interpolated models also enables adaptive model-size selection at inference time, improving the compute--accuracy trade-off for long-form reasoning tasks.

发表机构

  • Harvard University(哈佛大学)
  • IST Austria(奥地利科学技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑