arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HyperMCTS:用于长时程LLM智能体的超图增强MCTS

HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents

Tingsong Xiao, Nithish Balachandar Moudhgalya, Chandrayee Basu, Lichao Wang, Luyang Kong, Benjamin Z. Yao, Zhe Jiang, Jie Hao

arXiv 2609.33920首次发表:更新:

发表机构

University of Florida; Amazon(佛罗里达大学; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程LLM智能体搜索成本高且轨迹反馈复用不足的问题,提出无需训练的HyperMCTS方法,用跨轨迹超图增强MCTS树,通过超边聚合决策组回报指导选择,在DeepPlanning等任务上显著提升规划准确率并降低开销。

AI 中文摘要

长时程任务要求大语言模型(LLM)智能体在跨越整个解决方案的约束下协调决策。蒙特卡洛树搜索(MCTS)通过探索替代行动轨迹为测试时扩展提供了一种有前景的方法,但模型计算和环境交互使得搜索代价高昂。因此,高效搜索需要有效复用轨迹反馈。标准MCTS维护前缀特定的统计信息,而不显式累积跨不同路径重复出现的决策组的结果。为填补这一空白,我们提出HyperMCTS,一种无需训练的方法,用跨轨迹超图增强有序MCTS树。超边表示规范决策组,并在当前任务内累积其观测回报。我们的超图引导的HyperUCT选择规则将重叠超边的证据聚合成行动先验,使得在一个前缀下收集的结果能够为另一前缀下的选择提供信息,同时在树中保留执行历史。在DeepPlanning上,HyperMCTS相对于三个骨干模型各自的最强基线,将平均规划准确率提高了2.3至7.3个百分点。它使Qwen3.6-27B在Shopping Planning上超越Claude Opus 4.6(max),同时以更少的LLM调用和输出令牌数达到比所评估的基于MCTS的基线更高的准确率。SealQA实验进一步展示了在问答方面的改进。

英文摘要

Long-horizon tasks require large language model (LLM) agents to coordinate decisions under constraints that span an entire solution. Monte Carlo Tree Search (MCTS) offers a promising approach to test-time scaling by exploring alternative action trajectories, but model computation and environment interaction make search costly. Efficient search therefore requires effective reuse of trajectory feedback. Standard MCTS maintains prefix-specific statistics, without explicitly accumulating outcomes for decision groups that recur across different paths. To fill this gap, we propose HyperMCTS, a training-free method that augments an ordered MCTS tree with a cross-trajectory hypergraph. Hyperedges represent groups of canonical decisions and accumulate their observed returns within the current task. Our hypergraph-guided HyperUCT selection rule aggregates evidence from overlapping hyperedges into an action prior, allowing outcomes collected under one prefix to inform selection under another while preserving execution histories in the tree. On DeepPlanning, HyperMCTS improves average planning accuracy by 2.3--7.3 percentage points over the strongest baseline for each of three backbone models. It enables Qwen3.6-27B to outperform Claude Opus 4.6 (max) on Shopping Planning, while achieving higher accuracy with fewer LLM calls and output tokens than the evaluated MCTS-based baselines. SealQA experiments further demonstrate improvements in question answering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑