用于推理的思维级束搜索
Thought-Level Beam Search for Reasoning
- Princeton University(普林斯顿大学)
- MIT(麻省理工学院)
- Meta AI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大型推理模型测试时计算分配的低效问题,提出执行思维级束搜索的Gambit算法,在固定硬件预算下,其在准确率、吞吐量、token消耗等指标上均优于现有基线方法。
AI中文摘要:
测试时计算缩放是大型推理模型(LRMs)性能的主要驱动因素,但极端的低效率限制了现有方法,使得关键问题从“要投入多少”计算资源转向“将计算资源分配到哪里”。我们将测试时推理形式化为针对部分轨迹的受限计算分配问题。在固定硬件预算下,现有范式无法主动将计算资源分配给最有前景的部分进展:传统并行采样独立处理轨迹,会引发严重的内存瓶颈;而减法剪枝则会耗尽硬件资源,且无法主动、充分地调整输出分布。为克服这种两难局面,我们提出Gambit,一种执行思维级束搜索的推理算法。通过定期剪枝无前景的轨迹并从高质量前缀立即分支,Gambit通过探测隐藏状态的轻量评分器,动态将计算资源集中到最有前景的推理轨迹上,同时保持持续的高硬件利用率。在多个模型和基准上的大量评估表明,Gambit在现有基线中具有绝对优势。在相同硬件约束下,我们的方法在HMMT-24上比剪枝基线实现了高达6.7%的绝对准确率提升,在AIME-25上实现了3.3%的提升,在轨迹完成上的吞吐量提高了2倍以上,与标准并行采样相比,总token消耗最多减少68.5%。
英文摘要:
Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it. We formalize test-time reasoning as a constrained compute allocation problem over partial trajectories. Under a fixed hardware budget, existing paradigms fail to actively allocate the compute to the most promising partial progress: traditional parallel sampling treats traces independently and induces severe memory bottlenecks, while subtractive pruning starves hardware and fails to actively and sufficiently shift the output distribution. To overcome this dichotomy, we introduce Gambit, an inference algorithm that executes \emph{thought-level beam search}. By periodically pruning unpromising trajectories and immediately branching from high-quality prefixes, Gambit dynamically concentrates compute onto the most promising reasoning traces via a light-weight scorer probing hidden states while maintaining continuous high hardware utilization. Extensive evaluations across multiple models and benchmarks demonstrate that Gambit strictly dominates existing baselines. Under identical hardware constraints, our method yields up to a +6.7\% absolute accuracy gain on HMMT-24 and +3.3\% on AIME-25 over pruning baselines, delivers $>2\times$ higher throughput on trace completion, and reduces total token consumption by up to 68.5\% relative to standard parallel sampling.