arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应一致性图用于长时程智能体

Adaptive Consistency Graph for Long-Horizon Agents

Jiecong Wang, Hao Peng, Zhanyi Wang

arXiv 2609.32754首次发表:更新:

AI 中文总结

针对长时程智能体决策漂移问题,提出自适应一致性图(ACG),通过增量组织执行证据并构建需求中心视图,将GPT-5.6-luna平均成功率从44.5%提升至50.2%。

AI 中文摘要

大型语言模型智能体在短任务上通常能做出合理的局部决策,但当成功需要一系列依赖动作和工具调用的长序列时,其性能会下降。在执行过程中,任务要求、历史证据和当前执行状态可能逐渐脱节,导致后续决策偏离原始目标。我们通过引入自适应一致性图(ACG)来研究长时程执行问题。ACG在持久化图中增量组织执行证据及其来源,然后在有界上下文预算下为每个决策构建一个临时的、以需求为中心的视图。ACG不替换基础智能体的规划器或工具执行器,而是为每个决策提供结构化且可追踪的上下文视图。在匹配评估中,ACG将GPT-5.6-luna的平均成功率从ReAct的44.5%提升至50.2%,其中在BrowseComp-Plus上提升最大(73.5%对比62.4%)。我们进一步分析了轨迹结构和推理成本以表征这一改进。

英文摘要

Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, historical evidence, and the current execution state may gradually become disconnected, so later decisions can drift from the original objective. We study this problem by introducing the Adaptive Consistency Graph (ACG) for long-horizon execution. ACG incrementally organizes execution evidence and its provenance in a persistent graph, then constructs a temporary requirement-centered view for each decision under a bounded context budget. Rather than replacing the base agent's planner or tool executor, ACG provides a structured and traceable context view for each decision. In the matched evaluation, ACG improves GPT-5.6-luna's average success from 44.5\% with ReAct to 50.2\%, with the largest gain on BrowseComp-Plus (73.5\% versus 62.4\%). We further analyze trajectory structure and inference cost to characterize this improvement. Our code is available at https://github.com/yunsaijc/Adaptive-Consistency-Graph.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑