发表机构
University College London; National University of Singapore; Chinese Academy of Sciences; Brown University; University of Bristol; University of Edinburgh; Zhongguancun Academy; The Yangtze River Delta(伦敦大学学院; 新加坡国立大学; 中国科学院; 布朗大学; 布里斯托尔大学; 爱丁堡大学; 中关村科学院; 长三角)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM智能体现有轨迹级反思学习的缺陷,提出SKL方法,通过两种算法训练智能体学习有状态预测知识,在多任务上性能优于现有范式。
AI 中文摘要
随着大语言模型(LLM)智能体从经验中学习的能力不断提升,它们主要依赖轨迹级反思来提取洞见。从预测知识的视角来看,这种方法基于事件后的回溯而非预测性预见,会产生脆弱且依赖路径的启发式规则。为解决该问题,我们提出了有状态知识学习(SKL)。SKL将智能体的关注点从轨迹级总结转移到维护有状态知识:锚定状态的显式、声明式预测评估。我们首先展示了一个启发性示例,说明有状态知识如何提供粒度、增强泛化能力并实现知识自举。为进一步扩展该思路,我们通过自蒸馏(SKL-SD)和强化学习(SKL-RL)引入两种算法,训练智能体从经验中自主提取基于状态的预测知识并学习将其用于决策。在交互环境(WebShop、ScienceWorld)和复杂推理任务(ChessPuzzles)上的实验表明,为模型配备学习有状态预测知识的固有能力,显著优于当前基于反思的训练范式。
英文摘要
As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. Viewed through the lens of predictive knowledge, we argue that this approach operates on episodic hindsight rather than predictive foresight, yielding brittle, path-dependent heuristics. To address this, we propose Stateful Knowledge Learning (SKL). SKL shifts the agent's focus from trajectory-level summarization to maintaining Stateful Knowledge: explicit, declarative predictive assessments anchored to state. We first demonstrate a motivating example showing how stateful knowledge provides granularity, enhances generalization, and enables knowledge bootstrapping. To further scale up the idea, we introduce two algorithms via self-distillation (SKL-SD) and reinforcement learning (SKL-RL), training agents to autonomously extract state-grounded predictive knowledge from experience and learn to leverage it for policy making. Experiments on interactive environments (WebShop, ScienceWorld) and a complex reasoning task (ChessPuzzles) demonstrate that equipping models with the inherent ability to learn stateful predictive knowledge significantly outpaces current reflection-based training paradigms.