arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00888cs.LGcs.AI

匹配分布,而非计算量:训练后的多令牌预测头

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne, Yuan Gao, Tianwei Chen, George Zerveas, Ishmam Zabir, Xiren Zhou, Chris Quirk, Xia Song

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究在冻结推理模型上通过轻量级训练后处理实现多令牌预测加速,提出链感知验证松弛和自适应控制器,显著提升吞吐量并保持准确性。

中文摘要 AI 辅助

多令牌预测(MTP)通过使语言模型在每次前向传播中草拟多个后续令牌,提高了自回归生成的吞吐量,而对草拟令牌的验证步骤确保了骨干网络的令牌分布得以保留。每个开源的MTP系列发布(MiMo-7B、DeepSeek-V3、Qwen3)都在数十万亿令牌的完整预训练过程中与骨干网络联合训练其头部,从而在预训练时设定草拟质量。我们探究在冻结的推理模型上,对目标生成的思维链进行轻量级训练后处理是否足以达到相同的预期吞吐量加速,并研究如何优化基于此类检查点的服务时系统。我们提出三项发现。1) 在冻结的Qwen3-8B上,使用$K{=}3$个链式MTP头部,我们展示了在约$2.5$B令牌上使用简单交叉熵的训练后配方,在数学、编码和知识基准上达到或超过了联合训练的MiMo-7B的预期加速。我们的训练后配方相比MiMO-7B MTP基线的联合预训练,使用的MTP训练令牌减少了$10^3$-$10^4$倍。2) 我们提出了一种链感知的草拟令牌验证规则松弛,允许与骨干语言模型令牌分布的有界漂移。我们展示了这种松弛在每个基准上将预期加速提升$+12$至$+16\%$,同时保持任务准确性。3) 我们提出了一种自适应控制器,在推理时动态选择参与的MTP头部数量,并展示了使用固定最大MTP草拟长度可恢复高达$11$--$14\%$的加速损失。

英文摘要

Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present three findings. 1) On a frozen Qwen3-8B with $K{=}3$ chained MTP heads, we show that a post-training recipe with plain cross-entropy on $\approx\!2.5$B tokens reaches or exceeds the expected speedup of jointly trained MiMo-7B on math, coding and knowledge benchmarks. Our post-training recipe utilizes $10^3$-$10^4\times$ less MTP-training tokens as compared with joint pre-training of MiMO-7B MTP baseline. 2) We propose a chain-aware relaxation of draft token verification rule that allows a bounded drift from backbone language model token distribution. We show that this relaxation lifts expected speedups by $+12$ to $+16\%$ per benchmark while preserving task accuracy. 3) We propose an adaptive controller that dynamically chooses the number of MTP heads to be engaged at inference time and demonstrate recovery of upto $11$--$14\%$ loss in speedup using fixed maximum MTP draft length.

发表机构

  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

↑