arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01439cs.AI

DRelay:面向前缀感知并行推测解码修复的全局草稿上下文

DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair

Zhuoyu Wang, Junnan Huang, Xinyu Chen

首次发表
浏览论文内容

中文总结 AI 辅助

DRelay利用全局草稿上下文进行前缀感知的选择性修复,通过全局读取器和因果选择器纠正早期选择错误,延长接受前缀,在多个基准上提升推测解码的端到端性能。

中文摘要 AI 辅助

并行草稿生成降低了大型语言模型(LLMs)推测解码的草稿开销,但其收益仍受限于已接受前缀长度。即使候选池中存在正确令牌,单个早期选择错误也会阻止后续预测被使用。我们提出DRelay,它利用整个草稿块的全局信息,在目标模型验证之前对候选选择进行前缀感知的选择性修复。DRelay基于候选相关性和所选路径做出决策:一个全局读取器跨位置提取每个候选的预测信息。而一个因果选择器将全局读取提取的候选级信息与先前位置选择的令牌相结合,以确定当前位置的原生选择是否与全局证据和所选前缀一致。然后它决定保留或替换该令牌,从而修复早期错误并延长已接受前缀。我们进一步联合训练草稿主干和选择器,将候选支持学习与修复目标相结合,同时根据每个块位置对连续接受前缀的潜在贡献对修复损失进行加权。在H800 GPU上的八个多样化基准测试中,DRelay在平均接受长度和端到端解码性能上均持续优于DFlash、Domino和DSpark。在SGLang服务下,DRelay相对于DFlash、Domino和DSpark的平均端到端加速分别提高了14.7%-16.8%、8.7%-9.3%和8.1%-9.3%。

英文摘要

Parallel drafting reduces the drafting overhead of speculative decoding for large language models (LLMs), but its gains remain limited by the accepted prefix length. Even when the correct token is present in the candidate pool, a single early selection error prevents subsequent predictions from being used. We propose DRelay, which uses global information from the entire draft block to perform prefix-aware selective repair of candidate selections before target-model verification. DRelay bases its decisions on candidate correlations and the selected path: a global reader extracts predictive information across positions for each candidate. While a causal selector combines candidate-level information extracted by the global read with the tokens selected at preceding positions to determine whether the native choice at the current position is consistent with the global evidence and the selected prefix. It then decides whether to retain or replace the token, thereby repairing early errors and extending the accepted prefix. We further jointly train the draft backbone and the selector, combining candidate-support learning with a repair objective, while weighting the repair loss according to each block position's potential contribution to the consecutive accepted prefix. Across eight diverse benchmarks on an H800 GPU, DRelay consistently improves both average acceptance length and end-to-end decoding performance over DFlash, Domino, and DSpark. Under SGLang serving, DRelay improves average end-to-end speedup over DFlash, Domino, and DSpark by 14.7%-16.8%, 8.7%-9.3%, and 8.1%-9.3%, respectively.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

↑