指导的原则,行动的信息:通过知识抽象实现智能体进化
Principles that Guide, Actions that Inform: Agent Evolution via Knowledge Abstraction
浏览论文内容
中文总结 AI 辅助
针对LLM智能体经验泛化受限问题,提出SAGA框架,通过将交互轨迹抽象为原则并形成执行-抽象反馈循环,在ScienceWorld和ALFWorld上提升任务性能。
中文摘要 AI 辅助
大型语言模型(LLM)智能体在交互式环境中展现了强大的能力,但其从经验中持续进化的能力仍然有限。尽管微调能够实现适应,但它对参数访问的依赖和高昂的计算成本限制了其灵活性,尤其对于大规模和闭源的LLM而言。外部记忆提供了一种替代方案,允许智能体在不修改模型参数的情况下积累经验。然而,现有方法主要关注经验的表示和组织,而获得的知识仍然与特定任务和上下文紧密耦合,限制了泛化能力。一个关键挑战是如何将具体的交互转化为抽象且可复用的知识,以指导超越个体经验的未来决策。为应对这一挑战,我们提出了SAGA(通过经验基础抽象实现自我进化的智能体),一个用于LLM智能体中经验基础的知识抽象与利用的框架。SAGA逐步将交互轨迹转化为情景描述、可复用程序和具有明确适用条件的原则,同时保持与执行证据的联系。检索到的原则被实例化为任务特定的指导,并通过纠正性反馈和重采样来优化候选行动。这形成了一个执行-抽象反馈循环,其中积累的知识指导未来的交互,新经验持续更新分层记忆。在ScienceWorld和ALFWorld上的实验展示了任务性能的提升,消融研究强调了上下文实例化和行动调节对于利用原则级知识的重要性。
英文摘要
Large language model (LLM) agents have demonstrated strong capabilities in interactive environments, yet their ability to continually evolve from experience remains limited. Although fine-tuning enables adaptation, its dependence on parameter access and high computational costs restrict its flexibility, especially for large-scale and closed-source LLMs. External memory offers an alternative by allowing agents to accumulate experience without modifying model parameters. However, existing methods mainly focus on experience representation and organization, while the acquired knowledge remains tightly coupled with specific tasks and contexts, limiting generalization. A key challenge is how to transform concrete interactions into abstract and reusable knowledge that guides future decisions beyond individual experiences. To address this challenge, we propose SAGA (\underline{\textbf{S}}elf-evolving \underline{\textbf{A}}gents through Experience-\underline{\textbf{G}}rounded \underline{\textbf{A}}bstraction), a framework for experience-grounded knowledge abstraction and utilization in LLM agents. SAGA progressively transforms interaction trajectories into episodic descriptions, reusable procedures, and principles with explicit applicability conditions, while maintaining links to execution evidence. Retrieved principles are instantiated into task-specific guidance and used to refine candidate actions through corrective feedback and resampling. This creates an execution--abstraction feedback loop, where accumulated knowledge guides future interactions and new experiences continuously update hierarchical memory. Experiments on ScienceWorld and ALFWorld demonstrate improved task performance, with ablation studies highlighting the importance of contextual instantiation and action regulation for leveraging principle-level knowledge.
发表机构
- Alibaba Group(阿里巴巴集团)
- Shanghai Jiao Tong University(上海交通大学)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。