发表机构
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM推理时搜索的“推理盆地崩溃”问题,提出无需训练的结构感知选择方法BASIN,在24点游戏、MuSR上优于ToT,可提升推理时推理能力。
AI 中文摘要
大语言模型(LLMs)的推理时搜索常集中于一小部分结构或语义相似的轨迹,导致替代方案探索不足——这种失效模式我们称之为“推理盆地崩溃”。我们提出BASIN,一种无需训练的结构感知选择方法,它将推理状态分组为盆地并惩罚对同一策略的重复访问,从而在固定计算预算下重新分配搜索到真正不同的推理路径。在匹配的推理预算下,BASIN在24点游戏(Game of 24)上比思维树(Tree of Thoughts, ToT)提升高达22个百分点,在MuSR数据集上提升6.7个百分点。一种感知质量的变体QA-BASIN进一步提升鲁棒性,当无条件多样化过度探索时保留高质量盆地。为解释基于盆地的选择何时有效,我们引入冗余间隙Δ,它衡量搜索对正确与错误预测的集中程度差异:标准ToT通常在Δ≈0附近运行,而BASIN始终将Δ转为正值。更广泛地说,BASIN表明结构感知选择是提升推理时推理能力的简单通用方法。代码可在该https链接获取。
英文摘要
Inference-time search with large language models (LLMs) often concentrates on a small set of structurally or semantically similar trajectories, leaving alternative reasoning strategies underexplored---a failure mode we call \textit{reasoning basin collapse}. We introduce \textsc{BASIN}, a training-free, history-biased search method that groups reasoning states into basins and accumulates a revisit penalty on repeatedly selected basins, reallocating a fixed inference budget toward underexplored reasoning strategies. Under matched inference budgets, \textsc{BASIN} improves over Tree of Thoughts (ToT) by up to $+22$pp on Game of 24 and $+6.7$pp on MuSR. Because indiscriminate diversification can over-explore once search has found a promising basin, we further introduce \textsc{QA-BASIN}, a quality-aware variant that weakens the revisit penalty for high-quality basins and yields more robust gains. To characterize when basin-aware search helps, we introduce the \emph{redundancy gap} $Δ$, which measures the difference in search concentration between correct and incorrect predictions: standard ToT often operates near $Δ\approx 0$, whereas \textsc{BASIN} consistently shifts $Δ$ positive. Together, these results identify reasoning basin collapse as a failure mode of inference-time search and show that history-dependent bias provides a simple, training-free mechanism for escaping redundant reasoning under fixed compute. Code is available at https://github.com/GitHubLuCheng/basin
CommentsAccepted to NeurIPS'26