发表机构
Zhejiang University; Fudan University; Shanghai AI Laboratory; University of British Columbia; Stony Brook University; The Chinese University of Hong Kong(浙江大学; 复旦大学; 上海人工智能实验室; 不列颠哥伦比亚大学; 石溪大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本综述将大语言模型的推理视为搜索,系统化基于树搜索的推理进展,引入统一设计空间并倡导标准化计算报告,以优化推理权衡。
AI 中文摘要
随着预训练缩放定律趋近饱和,测试时缩放(Test-Time Scaling, TTS)已成为通过向固定模型分配推理时计算资源来提升推理能力的重要方向。从宏观层面看,TTS将推理重新表述为在部分推理状态空间上的搜索。尽管思维链(Chain-of-Thought, CoT)揭示了中间步骤,但常见实例依赖单轨迹解码,限制了从早期错误中恢复和探索的能力。本综述系统化了基于树搜索的推理的最新进展,将推理视为特定实例的优化而非解码。我们追溯了从无信息搜索到蒙特卡洛树搜索(Monte Carlo Tree Search, MCTS)的演变,强调基于采样的控制如何支持原则性的探索-利用权衡。为统一碎片化的文献,我们引入了涵盖搜索拓扑、评估信号和控制动态的统一设计空间,并倡导标准化的计算报告抽象,以明确且可比较计算-准确率权衡。
英文摘要
As pretraining scaling laws approach saturation, Test-Time Scaling (TTS) has emerged as an important direction for improving reasoning by allocating inference-time compute to a fixed model prior. Viewed at a high level, TTS reframes inference as search over a space of partial reasoning states. While Chain-of-Thought (CoT) exposes intermediate steps, common instantiations rely on single-trajectory decoding, limiting recovery from early errors and exploration. This survey systematizes recent progress in tree-search-based reasoning, viewing inference as instance-specific optimization rather than decoding. We trace the evolution from uninformed search to Monte Carlo Tree Search (MCTS), highlighting how sampling-based control supports principled exploration-exploitation trade-offs. To unify a fragmented literature, we introduce a Unified Design Space spanning search topology, evaluation signals, and control dynamics, and advocate a standardized compute-reporting abstraction to make compute-accuracy trade-offs explicit and comparable.
CommentsAccepted by EMNLP'2026