arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReTrace:用于投机解码的拒绝轨迹条件方法

ReTrace: Rejected-Trajectory Conditioning for Speculative Decoding

Luxi Lin, Zhanpeng Zeng, Shuang Peng, Songwei Liu, Rongrong Ji

arXiv 2608.29748首次发表:更新:

发表机构

Xiamen University(厦门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ReTrace方法,针对投机解码中被拒绝后缀的有效信息,通过跨轮条件保留优化拒绝后缀,无需额外计算即可提升Qwen3模型在多任务上的解码效率,且可与现有方法结合。

AI 中文摘要

投机解码通过轻量级草稿模型提出多个候选token,再由更大的目标模型并行验证,以此加速自回归语言模型推理。但在首次拒绝后,标准的基于前缀的验证会丢弃剩余的草稿后缀,生成和验证这些位置的计算无法为解码进度做贡献。以DFlash为研究对象,我们发现被拒绝后缀中的被拒绝位置仍可能与目标的后续内容对齐,说明草稿模型尽管存在局部token级错误,仍能保留有用的语义和结构信息。受此观察启发,并借鉴条件扩散的思路,我们提出了ReTrace——一种拒绝轨迹条件方法,该方法让每个草稿块以上一轮的拒绝后缀为条件,而非仅从新的掩码占位符生成。ReTrace保留拒绝后缀的隐藏表示,将其与下一个草稿块对齐,利用同一验证过程中来自目标感知的校正信号对其进行优化,并通过门控残差融合将其纳入草稿模型的输入嵌入。由于被拒绝token从未被提交,且目标侧的验证保持不变,ReTrace保留了投机解码的无损属性,无需额外的模型前向传播。针对Qwen3模型在数学推理、代码生成和开放式对话上的实验表明,ReTrace相比其DFlash骨干,能持续提升平均接受长度和端到端解码速度。通过引入跨轮条件而不修改轮内提案生成,ReTrace与现有的草稿改进方法基本正交,可与其结合以获得进一步提升。

英文摘要

Speculative decoding accelerates autoregressive language model inference by having a lightweight draft model propose multiple candidate tokens, which are then verified in parallel by a larger target model. However, after the first rejection, standard prefix-based verification discards the remaining draft suffix, so the computation spent generating and verifying those positions does not contribute to decoding progress. Focusing on DFlash, we show that rejected positions in a rejected suffix may still align with the target continuation, indicating that the draft model can retain useful semantic and structural information despite local token-level errors. Motivated by this observation and inspired by conditional diffusion, we introduce ReTrace, a rejected-trajectory conditioning method that conditions each draft block on the rejected suffix from the previous round rather than generating it from fresh mask placeholders alone. ReTrace retains the hidden representations of the rejected suffixes, aligns them with the next draft block, refines them using target-aware correction signals from the same verification pass, and admits them into the drafter's input embeddings through gated residual fusion. Because rejected tokens are never committed and target-side verification remains unchanged, ReTrace preserves the lossless property of speculative decoding without requiring an additional model forward pass. Experiments with Qwen3 models across mathematical reasoning, code generation, and open-ended dialogue demonstrate that ReTrace consistently improves average acceptance length and end-to-end decoding speed over its DFlash backbone. By introducing cross-round conditioning without modifying within-round proposal generation, ReTrace is largely orthogonal to existing drafting improvements and might be combined with them for further gains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑