arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoMeRL:通过降阶效用状态平衡自进化智能体记忆中的反馈覆盖与记忆-奖励陷阱

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

Yi Yang, Zhennan Chen, Yihong Zhuang, Tiehan Fan, Yinan Chen, Jian Li, Jian Yang, Ying Tai

arXiv 2608.02508首次发表:更新:

发表机构

Nanjing University; Xiamen University; Zhejiang University(南京大学; 厦门大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RoMeRL通过降阶效用状态解决自进化LLM智能体记忆的反馈分散与记忆-奖励陷阱问题,在ALFWorld等基准上显著提升任务性能、降低记忆规模与LLM调用

AI 中文摘要

基于学习的自进化大语言模型(LLM)智能体记忆系统面临两个紧密关联的挑战:其一,按轨迹索引的效用会随交互历史增长,导致有限的反馈分散到不断扩张的状态空间中;其二,由于轨迹级奖励会共同分配给协同检索的记忆,无关经验可能会收到误导性的效用更新,进而陷入记忆-奖励陷阱。为应对这些挑战,我们提出了降阶记忆强化学习(RoMeRL),该方法通过按结果极性和记忆动力学分解的固定维度单任务记忆状态,表征不断增长的轨迹索引效用空间。RoMeRL通过一组固定的语义坐标整合新经验,其内容会随时间更新或替换,从而将反馈集中在有界的效用支撑集上。理论上,我们证明这种降阶参数化可提升每个效用坐标收到的平均反馈量,并在通用坐标转移模型下刻画错误坐标的稳态占用情况。实验中,在ALFWorld和LifelongAgentBench两个基准上,RoMeRL提升了任务性能,将冷Q(Cold-Q)比例降低80.0%,反馈密度提升约6.0倍,维持的记忆大小减少84.4%,并减少21.1%的大语言模型调用。这些结果表明,降阶效用状态可支撑高效的自进化智能体记忆,同时限制持续的奖励污染。代码可访问:this https URL

英文摘要

Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory-reward trap. To address these challenges, we introduce Reduced-Order Memory Reinforcement Learning (RoMeRL), which represents the growing trajectory-indexed utility space using a fixed-dimensional per-task memory state factorized by outcome polarity and memory dynamics. RoMeRL incorporates new experiences through a fixed set of semantic coordinates whose contents are updated or replaced over time, thereby concentrating feedback over a bounded utility support. Theoretically, we show that this reduced-order parameterization increases the average feedback received by each utility coordinate and characterize the steady-state occupancy of erroneous coordinates under a generic coordinate-transition model. Empirically, across ALFWorld and LifelongAgentBench, RoMeRL improves task performance, reduces the Cold-Q ratio by 80.0%, increases feedback density by approximately 6.0 times, reduces the maintained memory size by 84.4%, and cuts LLM calls by 21.1%. These results show that reduced-order utility states support efficient self-evolving agent memory while limiting persistent reward contamination. Code is available at: https://github.com/YOUNG-fnxm/RoMeRL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑