arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

程序图:面向LLM智能体的自演化执行结构

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

Yuxing Lu, Yicheng Chen, Shanchan Wu, Sercan Ö. Arık

arXiv 2609.09153首次发表:更新:

发表机构

Google; Georgia Institute of Technology; Peking University(谷歌公司; 佐治亚理工学院; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出程序图框架,将程序性知识组织为三元组图,通过引导模型提供步骤级情境引导并自演化编辑图结构,在多个数据集上超越记忆基线,提升LLM智能体长时程决策性能。

AI 中文摘要

大型语言模型越来越多地被部署为在长时程范围内规划并通过外部工具行动的智能体。大多数智能体通过对不断累积的历史进行无约束生成来选择动作,这使得关于做什么、按什么顺序做以及在何种条件下做的程序性知识保持隐式。随着轨迹变长,智能体可能失去对目标的追踪、无序地调用工具,并重复无成效的动作。我们引入了程序图:正如知识图谱将事实性知识组织为(实体,关系,实体)三元组以回答“是什么”的问题,程序图将程序性知识组织为(程序,关系,程序)三元组以回答“该做什么”的问题。在每个决策步骤,该框架定位智能体的活动节点,一个引导模型将周围的子图转换为步骤级的情境引导,该引导影响求解器的下一个动作但不强制规定它。该图是自演化的:一个LLM精炼器将失败的轨迹与成功的轨迹进行对比,并编辑图的拓扑和属性,提交那些保持或改进留出验证性能的编辑,同时保留被拒绝的编辑以阻止重复。从一个最小的骨架开始,该循环构建出与手工设计的图相匹配或超越它们的图。它还能修复一个有缺陷的专家先验。在多个数据集、任务类型和LLM上,程序图相对于基于记忆的基线带来了一致的改进,并且自演化进一步提升了性能,而无需手动工程。

英文摘要

Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order, and under which conditions. As trajectories lengthen, agents can lose track of their objectives, invoke tools out of order, and repeat unproductive actions. We introduce the Procedural Graph: just as a knowledge graph organizes factual knowledge into (entity, relation, entity) triplets for what-is questions, a Procedural Graph organizes procedural knowledge into (procedure, relation, procedure) triplets for what-to-do questions. At each decision step, the framework localizes the agent's active node, and a guidance model translates the surrounding subgraph into step-level situational guidance that biases the solver's next action without dictating it. The graph is self-evolving: an LLM refiner contrasts failed trajectories with successful ones and edits the graph's topology and attributes, committing edits that preserve or improve held-out validation performance while retaining rejected ones to discourage repetition. Starting from a minimal skeleton, the loop builds graphs that match or surpass hand-designed ones. It can also repair a flawed expert prior. Across multiple datasets, task types, and LLMs, the Procedural Graph delivers consistent gains over memory-based baselines, and self-evolution further improves performance without manual engineering.

Comments36 pages including references and appendices, 6 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑