发表机构
Korea University; POSTECH(高丽大学; 浦项科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SpecScale服务系统,通过早期剪枝、去重计算和延迟验证三种技术,解决投机执行中搜索空间爆炸和细粒度验证的挑战,显著提升LLM推理吞吐量和延迟,同时保持答案质量。
AI 中文摘要
测试时扩展最近成为一种强大的方法,通过在推理期间分配额外计算来改善LLM推理,显著提高数学和编码等挑战性任务的准确性。为了加速推理路径的探索,最近的研究提出了投机执行。然而,我们表明支持投机执行给LLM服务系统带来了两个独特的挑战:(1)候选路径搜索空间的爆炸性增长,(2)对候选路径进行频繁、细粒度的验证任务。为了解决这些挑战,本文提出了SpecScale,一个用于高效投机执行的服务系统。我们引入了三种技术来调和延迟与计算开销之间的权衡:(1)早期剪枝低质量候选路径,(2)对冗余候选路径去重计算,(3)延迟细粒度验证任务。我们在具有挑战性的推理基准上评估SpecScale,包括MATH和Olympiad。我们的结果表明,SpecScale显著优于非投机和最近的投机方法,在保持答案质量的同时,在吞吐量和延迟方面带来了实质性改进。
英文摘要
Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, substantially enhancing accuracy on challenging tasks such as mathematics and coding. To accelerate the exploration of reasoning paths, recent studies proposed speculative execution. However, we show that supporting speculative execution poses two unique challenges for LLM serving systems: (1) an explosion in the search space of candidate paths and (2) frequent, fine-grained verification tasks for candidates. To address these challenges, this paper proposes SpecScale, a serving system for efficient speculative execution. We introduce three techniques to reconcile the trade-off between latency and computational overhead: (1) early pruning of low-quality candidate paths, (2) deduplicating computation across redundant candidate paths, and (3) deferring fine-grained verification tasks. We evaluate SpecScale on challenging reasoning benchmarks, including MATH and Olympiad. Our results show that SpecScale significantly outperforms both non-speculative and recent speculative approaches, delivering substantial improvements in throughput and latency while preserving answer quality.
Comments14 pages