可证明的测试时扩展:LLM推理中的束搜索
Provable Test-Time Scaling for Beam Search in LLM Reasoning
浏览论文内容
中文总结 AI 辅助
本文研究LLM推理中束搜索的测试时计算保证,提出置信度过滤束搜索(CF-Beam),将覆盖依赖性从二次降至近线性,并证明其优于Best-of-N等序列级方法,实验验证其在困难实例和长时域下更稳健。
中文摘要 AI 辅助
基于束搜索的测试时方法通过早期剪除无效推理路径,为改善大语言模型(LLM)在长时程生成任务上的性能提供了一种有效途径,从而显著提升推理效率并带来更有利的测试时成本扩展。尽管经验上取得了巨大成功,但束搜索的理论理解仍然有限。本文研究了一种常用束搜索框架的测试时计算保证,该框架利用模型内部的对数似然进行中间评分,仅在完整响应生成后依赖外部奖励模型。我们首先为朴素束搜索建立了下界,表明最优响应存活至少需要 $\Omega(C^\star(x)^2)$ 个样本,其中 $C^\star(x)$ 是提示 $x$ 的令牌级覆盖系数。这促使我们提出了改进的置信度过滤束搜索(CF-Beam),在固定时域、间隔和目标精度下,该算法将充分覆盖依赖性从二次降低到近乎线性(在前缀竞争性条件下)。随后我们证明,CF-Beam 的遗憾值上界由罕见失败事件的概率和由路径级覆盖系数缩放的奖励估计误差组成,其中罕见失败项随每步采样增加而消失。我们的结果凸显了束搜索相对于序列级推理方法(如 Best-of-N 和 Best-of-Majority)的根本优势。这些方法的保证通常涉及随时域 $L$ 指数增长的覆盖系数,而 CF-Beam 通过一个随 $L$ 多项式增长的令牌级覆盖系数来控制主导的搜索诱导项。我们的数值实验进一步证实,束搜索在困难实例和更长的推理时域下更为稳健。
英文摘要
Beam-search-based test-time methods provide an effective way to improve large language model (LLM) performance on long-horizon generation by pruning invalid reasoning paths early, leading to significantly improved reasoning efficiency and more favorable test-time cost scaling. Despite strong empirical success, the theoretical understanding of beam search remains limited. In this paper, we study the test-time compute guarantee of the commonly used beam search framework that uses the model's internal log-likelihood for intermediate scoring, while relying on an external reward model only after a complete response is generated. We first establish a lower bound for vanilla beam search, showing that at least $Ω(C^\star(x)^2)$ samples are required for the optimal response to survive, where $C^\star(x)$ is the token-level coverage coefficient for prompt $x$. This motivates our modified confidence-filtered beam search (CF-Beam), which reduces the sufficient coverage dependence from quadratic to nearly linear under prefix competitiveness, for fixed horizon, gap, and target accuracy. We then show that the regret of CF-Beam is upper-bounded by the probability of rare failure events and the reward estimation error scaled by a path-level coverage coefficient, where the rare-failure term vanishes as per-step sampling increases. Our results highlight a fundamental advantage of beam search over sequence-level inference methods such as Best-of-N and Best-of-Majority. While the guarantees of these approaches typically involve coverage coefficients that grow exponentially with the horizon $L$, CF-Beam controls the dominant search-induced term through a token-level coverage coefficient that scales polynomially with $L$. Our numerical experiments further confirm that beam search is more robust on hard instances and under increasing reasoning horizons.
发表机构
- The Ohio State University(俄亥俄州立大学)
- University of Pennsylvania(宾夕法尼亚大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。