AI 中文总结
针对并行推测解码的长序列精度下降问题,提出预算感知动态修复树CURE,通过定位高不确定性token修复错误,在多个基准上提升平均接受长度并实现2.66-3.49倍端到端速度提升。
AI 中文摘要
推测解码通过将草稿生成与目标验证相结合,缓解了自回归大语言模型(LLMs)中顺序生成的延迟问题。然而,现有的并行草稿后端往往在长序列生成时出现快速的精度下降,导致验证阶段的拒绝率较高,最终的时钟速度提升效果不佳。我们发现,草稿错误并非均匀分布,而是通常源于局部的高不确定性token,这些token会破坏后续的生成轨迹。基于这种token错误模式,我们提出了CURE,这是一种预算感知的动态修复树,旨在在不确定性焦点处修复错误,同时不会产生过高的树验证开销。具体而言,我们的方法利用预测置信度边界在块并行草稿中动态定位候选错误token,仅在这些脆弱节点处扩展有界修复路径,并采用一种新颖的修复重同步机制,在验证后重新对齐草稿状态。在代码生成基准(HumanEval、MBPP和LiveCodeBench-lite)以及数学推理基准(GSM8K)上的评估表明,与未进行修复的并行基线相比,CURE将平均接受长度提升了4.2%-7.5%,对应于仅目标解码的端到端速度提升达到2.66-3.49倍。此外,我们提供了一个即插即用的修复模块,可与标准并行草稿框架兼容,同时我们还分析了草稿计算与验证效率之间的权衡关系。
英文摘要
Speculative decoding mitigates the latency of sequential generation in autoregressive Large Language Models (LLMs) by interleaving draft generation with target verification. However, existing parallel drafting backends often suffer from rapid accuracy degradation over long horizons, leading to high rejection rates during verification and suboptimal wall-clock speedups. We observe that drafting errors are not uniformly distributed but typically stem from localized high-uncertainty tokens that destabilize downstream generation trajectories. Motivated by this token error pattern, we propose CURE, a budget-aware dynamic repair tree designed to repair errors at uncertainty focal points without incurring prohibitive tree-verification overheads. Specifically, our method uses predictive confidence margins to dynamically locate candidate error tokens within a block-parallel draft, expands bounded repair paths only at these fragile nodes, and employs a novel repair resynchronization mechanism to realign draft states post-verification. Evaluations on code-generation benchmarks (HumanEval, MBPP, and LiveCodeBench-lite) and mathematical reasoning benchmark (GSM8K) demonstrate that CURE increases the average accepted length by 4.2-7.5% over parallel baselines without repair, translating to an end-to-end speedup of $2.66-3.49\times$ over target-only decoding. Furthermore, we provide a plug-and-play repair module compatible with standard parallel drafting frameworks. We also characterize the trade-off between draft compute and verification efficiency.
Comments9 pages, 2 figures, 5 tables