arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10676cs.AI

基于树状记忆的自修正长程搜索智能体

Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Aijun Yang, Qianxue Guo, Ziyi Huang, Yuxuan Chen, Shiyou Qian, Jian Cao

AI总结:

提出ReTree树状记忆机制,解决LLM搜索智能体的上下文增长与错误推理问题,在四个基准上优于Full-Trajectory ReAct,准确率最高提升25.6个百分点。

AI中文摘要:

基于大语言模型(LLM)的搜索智能体通过与外部环境的多步交互回答问题,但向LLM提供完整的执行轨迹会导致无界的上下文增长并引入噪声。现有压缩方法以丢失重要细节为代价减少上下文,且常替换错误事实却不修复由此产生的下游推理。为解决该问题,我们提出ReTree,一种面向搜索智能体的自修正树状记忆机制。ReTree构建了有界的逐步推理上下文,同时保留与源关联的证据。它将搜索建模为证据树,节点存储有界摘要、证据和修正历史。当新检索的证据与早期主张矛盾时,ReTree回溯到引入该主张的节点,替换过时证据、重新生成摘要、修剪受影响分支并恢复搜索。基于源的证据溯源支持可靠的冲突定位,并使最终主张可追溯到检索的段落。在四个公开的问答与搜索基准上的实验表明,ReTree的表现始终优于Full-Trajectory ReAct,答案准确率最高提升25.6个百分点(pp);Full-Trajectory ReAct的平均最大逐步推理上下文是ReTree的1.27至1.51倍。这些结果确立了ReTree作为长程搜索的有效自修正记忆抽象的地位。

英文摘要:

Large language model (LLM)-based search agents answer questions through multi-step interactions with external environments. However, providing complete execution trajectories to the LLM causes unbounded context growth and introduces noise. Existing compression methods reduce context at the cost of important details and often replace erroneous facts without repairing downstream reasoning derived from them. To address this problem, we propose ReTree, a self-correcting tree-structured memory mechanism for search agents. ReTree constructs a bounded per-step reasoning context while preserving source-linked evidence. It models search as an evidence tree whose nodes store bounded summaries, evidence, and revision histories. When newly retrieved evidence contradicts an earlier claim, ReTree traces back to the node where the claim was introduced, replaces outdated evidence, regenerates summaries, prunes affected branches, and resumes search. Source-grounded evidence provenance supports reliable conflict localization and keeps final claims traceable to retrieved passages. Experiments on four public question-answering and search benchmarks show that ReTree consistently outperforms Full-Trajectory ReAct, improving answer accuracy by up to 25.6 percentage points (pp); the average maximum per-step reasoning context of Full-Trajectory ReAct is $1.27$--$1.51\times$ that of ReTree. These results establish ReTree as an effective self-correcting memory abstraction for long-horizon search.

↑