arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Osprey:目标无关的预训练使推测解码中的草稿模型更强大

Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

Fengxiang Bie, Yuqing Jian, Yifan Yu, Zhongzhu Zhou, Zelei Shao, Ben Athiwaratkun, Shuaiwen Leon Song, Chenfeng Xu, Xiaoxia Wu, Tianyi Zhang

arXiv 2609.09338首次发表:更新:

发表机构

Together AI; The University of Sydney; University of Illinois Urbana-Champaign; The University of Texas at Austin(Together AI; 悉尼大学; 伊利诺伊大学厄巴纳-香槟分校; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Osprey通过目标无关的预训练和轻量级适配,使草稿模型跨目标迁移,提升推测解码的平均接受长度和推理速度。

AI 中文摘要

推测解码对于加速大型语言模型推理至关重要。然而,其加速效果是脆弱的:草稿模型通常针对单一目标模型的狭窄分布进行训练,在工作负载变化时其接受率会急剧下降。这与现代大型语言模型的发展形成了鲜明对比,在后者中,目标模型之所以受到重视,正是因为它们通过大规模预训练获得了广泛的泛化能力。我们认为,自然的补救措施——预训练——难以应用于草稿模型,因为现有的方法都是目标特定的:草稿模型消耗目标模型的隐藏状态,并在目标模型的逻辑上进行蒸馏,因此预训练必须针对每个目标重复进行。我们提出了Osprey,它从现成的预训练小语言模型中引导出草稿模型,将广泛的预训练视为可重用的、目标无关的资产,并将每个目标的适配工作减少为轻量级的适应步骤。实现这一点需要克服两个挑战:小语言模型比延迟受限的草稿模型所能承受的深度要深得多,并且其预训练的计算能力必须保持完整,同时草稿模型需要学会接收目标模型的隐藏状态并以目标模型的词汇表生成标记。Osprey通过剪枝到浅层骨干网络、使用目标无关的下一个标记预训练恢复其语言建模能力,并通过词汇对齐、零初始化的QKV扩展以及从目标模型输出分布进行蒸馏来适应每个目标,从而解决了这两个问题。实验上,一个预训练的Osprey骨干网络可跨目标迁移,并使Qwen3-8B的平均接受长度提高16.1%,Llama-3.3-70B-Instruct提高21.2%,229B MiniMax-M2.5提高22.7%(每秒标记数提高17.5%),在域外和多语言数据上收益最大。我们的代码可在https://this https URL获取。

英文摘要

Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models are valued precisely for the broad generalization they acquire through large-scale pretraining. We argue that the natural remedy, pretraining, has been hard to apply to drafters because existing recipes are target-specific: the drafter consumes the target's hidden states and is distilled on the target's logits, so pretraining must be repeated for each target. We introduce Osprey, which instead bootstraps drafters from off-the-shelf pretrained small language models, treating broad pretraining as a reusable, target-agnostic asset and reducing per-target work to a lightweight adaptation step. Realizing this requires overcoming two challenges: small LMs are far deeper than a latency-bound drafter can afford, and their pretrained computation must remain intact while the drafter learns to ingest target hidden states and emit tokens in the target's vocabulary. Osprey addresses both by pruning to a shallow backbone, restoring its language-modeling capability with target-agnostic next-token pretraining, and adapting it to each target through vocabulary alignment, zero-initialized QKV expansion, and distillation from the target model's output distribution. Empirically, a single pretrained Osprey backbone transfers across targets and improves mean acceptance length by 16.1% for Qwen3-8B, 21.2% for Llama-3.3-70B-Instruct, and 22.7% for the 229B MiniMax-M2.5 (with 17.5% higher tokens per second), with the largest gains on out-of-domain and multilingual data. Our code is available at https://github.com/LeanModels/Osprey.

CommentsAccepted at EMNLP 2026. 21 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑