发表机构
University of Science and Technology of China; Singapore Management University; National University of Singapore(中国科学技术大学; 新加坡管理大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Knowledge Weaver强化学习框架,利用互信息反馈和边际成功奖励训练语言模型策展智能体轨迹知识,在ALFWorld和WebShop上分别提升成功率16.9和18.7个百分点。
AI 中文摘要
大型语言模型(LLM)智能体可以通过重用从过去交互中提炼的知识来提高其性能。然而,将新经验整理成一个随着增长而变得更有用的知识库仍然具有挑战性。有效的知识积累应限制条目之间的冗余重叠,并确保新知识在库已提供的内容之外做出贡献。然而,使用组相对策略优化(GRPO)在独立任务成功率上训练策展者可能会强化通用指导,即使它与现有知识重复。因此,我们提出了Knowledge Weaver,一个强化学习框架,训练语言模型从智能体轨迹中策展可重用的知识。我们将受令牌级互信息(MI)启发的反馈与边际成功奖励相结合,以指导知识积累。这些信号共同鼓励策展者保留经验中的独特信息,并生成在添加到现有知识时能提高任务成功率的条目。独立成功奖励也倾向于那些本身有用的条目。在ALFWorld和WebShop上,Knowledge Weaver在k=10个检索条目下实现了54.0%和42.0%的平均成功率,分别超过GRPO 16.9和18.7个百分点。其知识库在冻结执行器的情况下,在ALFWorld整体成功率和WebShop得分上也优于评估的基于提示的库和已建立的库,包括人工编写的库。我们的代码库可在以下网址获取:https://this URL。
英文摘要
Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions. However, curating new experiences into a knowledge bank that becomes more useful as it grows remains challenging. Effective knowledge accumulation should limit redundant overlap among entries and ensure that new knowledge contributes beyond what the bank already provides. Yet training a curator with Group Relative Policy Optimization (GRPO) on standalone task success can reinforce general guidance even when it duplicates existing knowledge. Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories. We couple feedback inspired by token-wise mutual information (MI) with marginal success rewards to guide knowledge accumulation. Together, these signals encourage the curator to preserve distinct information from experience and produce entries that improve task success when added to existing knowledge. Standalone success rewards also favor entries that are useful on their own. On ALFWorld and WebShop, Knowledge Weaver achieves mean success rates of 54.0\% and 42.0\% with k=10 retrieved entries, exceeding GRPO by 16.9 and 18.7 percentage points, respectively. Its knowledge banks also outperform the evaluated prompt-based and established banks, including human-written banks, in overall ALFWorld success rate and WebShop score with the executor frozen. Our codebase is available at https://github.com/LaoKuiZe/Knowledge-Weaver.