arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAST:来自单次前向块草稿器的成本感知推测树

CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters

Jungseob Lee, Sugyeong Eo

arXiv 2610.00321首次发表:更新:

发表机构

Korea University; Yonsei University Mirae Campus(高丽大学; 延世大学未来校区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CAST通过将块草稿器的候选打包成树并在单次目标传递中验证,依据延迟测量自适应树宽,在多种部署中最高提速43%,且不改变输出分布。

AI 中文摘要

推测解码通过廉价地草拟未来令牌并与目标模型并行验证它们,从而加速大型语言模型的推理。块草稿器在一次前向传递中对整个未来令牌块进行评分,但标准解码仅验证得分最高的链并丢弃其他候选。由于这些候选已经被评分,验证更多的候选会增加目标计算量,但不会增加额外的草拟成本。我们引入了CAST(成本感知推测树),它将这些候选打包成一棵树,并在单次目标传递中验证它,同时保持目标模型、草稿器权重和解码规则不变。为了决定树的宽度,CAST在下一个候选的预期收益超过其增加的验证时间时添加候选。因此,宽度根据延迟测量适应每次部署,而无需扫描各种宽度。我们在三个GPU代和两个模型家族的五个领域上评估了CAST。在其预测宽度下,CAST在所有八种设置中都比标准链更快,最高提升达43%。我们还发现最佳宽度强烈依赖于部署环境。在验证成本在核边界处跳跃的情况下,128令牌树仅比标准链快2%,而预测宽度下的树则快20%。此外,我们证明了CAST在贪婪解码和采样解码下都保持目标输出分布不变。代码可在该https URL获取。

英文摘要

Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding verifies only the top-scoring chain and discards the other candidates. Because these candidates are already scored, verifying more of them adds target computation but no extra drafting. We introduce CAST (Cost-Aware Speculative Trees), which packs these candidates into a tree and verifies it in a single target pass, leaving the target model, drafter weights, and decoding rule untouched. To decide how wide the tree should be, CAST adds candidates while the expected gain from the next one outweighs the verification time it adds. The width therefore adapts to each deployment from a latency measurement, without sweeping over widths. We evaluate CAST across five domains on three GPU generations and two model families. At its predicted width, CAST is faster than the standard chain in all eight settings, by up to 43%. We also find that the best width depends strongly on the deployment. Where verification cost jumps at a kernel boundary, a 128-token tree is only 2% faster than the standard chain, whereas the tree at the predicted width is 20% faster. Furthermore, we prove that CAST leaves the target output distribution unchanged under both greedy and sampled decoding. Code is available at https://github.com/js-lee-AI/CAST.

Comments28 pages, 7 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑