arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

恢复投机解码的离策略监督

Recovering Off-Policy Supervision for Speculative Decoding

Jungseob Lee, Chanjun Park, Sugyeong Eo, Hyeonseok Moon

arXiv 2609.38795首次发表:更新:

发表机构

Korea University; Soongsil University; Yonsei University Mirae Campus; Sookmyung Women’s University(高丽大学; 崇实大学; 延世大学未来校区; 淑明女子大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于展开的框架,通过锚标签重标记和展开内锚恢复投机解码草稿器的完整监督,在固定语料库上显著提升接受长度并节省计算。

AI 中文摘要

投机解码的块草稿器通常是在外部模型编写的语料库上训练的,其中单个离策略令牌会使块中所有后续槽位的监督失效。现有方法丢弃这些分歧槽位,导致严重的监督损失。为解决此问题同时保留训练语料库,我们提出了一种基于展开的训练框架,通过两个互补组件恢复完整监督。第一个组件,锚标签重标记(ALR),用贪婪目标展开的分布替换语料库标签,恢复所有预测槽位的有效监督。第二个组件,展开内锚(IRA),将草稿块直接放置在展开内,使草稿器暴露于目标生成的上下文,重用预计算的展开特征而无需额外目标成本。在固定的视觉-语言和文本语料库上,我们的框架将贪婪接受长度比DFlash提高多达36.5%,并持续优于擦除基线。值得注意的是,我们方法的一个训练周期就超过了最佳的擦除调度。经过三个周期后,它匹配了在目标重新生成的响应上训练的接受长度。这些结果表明,我们的框架提供了一种有效且计算高效的方法,用于在固定语料库上训练投机草稿器,而无需修改原始文本。代码可在此https URL获取。

英文摘要

Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the drafter to target-generated context, reusing precomputed rollout features at no additional target cost. Across fixed vision-language and text corpora, our framework increases greedy accepted length by up to 36.5% over DFlash and consistently outperforms erasing baselines. Notably, a single epoch of our method surpasses the best erase schedules. After three epochs, it matches the acceptance length of training on target-regenerated responses. These results show that our framework provides an effective and compute-efficient approach for training speculative drafters on fixed corpora without modifying the original text. Code is available at https://github.com/js-lee-AI/ALR-IRA.

Comments22 pages, 4 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑