Revisiting Tree Search for LLMs: Gumbel and Sequential Halving for Budget-Scalable Reasoning
重新审视用于大语言模型的树搜索:Gumbel与顺序削减用于可扩展的推理
专题命中 模型式强化学习 :model-based reinforcement learning(abstract);分类 cs.AI、cs.LG
AI总结 本文提出ReSCALE,通过将Gumbel采样和顺序削减替代传统方法,实现大语言模型在推理中的可扩展推理能力提升,解决了搜索预算增加时准确率下降的问题。
Comments The paper has been accepted to the ICAPS-2026 conference. 5 pages, 2 figures